How Anthropic's Claude made a math breakthrough during its unsuccessful 54-hour attempt to solve the Riemann hypothesis, aided by an employee's encouragement
The world's smartest AI models are now superhuman at math. They still respond to moral support and encouragement from mere humans.
Context & Ripple Effects
Anthropic had already disclosed that an unreleased Claude model made unexpected progress on a related problem while failing at the Riemann hypothesis. The latest account adds a workflow detail: an employee's encouragement accompanied the model's extended effort, following Anthropic's earlier account of the unsuccessful Riemann attempt.
The episode lands as mathematicians are testing AI against unpublished research questions through the First Proof experiment, while AI researchers increasingly treat mathematics as a measure of reasoning-model progress. It therefore offers evidence about both model capability and the human supervision around long-running research tasks.
First-order effects
- Anthropic gains a concrete demonstration that Claude can generate useful mathematical progress even when its top-level research objective is not completed.
- The reported role of employee encouragement makes human interaction part of the documented working process around Claude's lengthy mathematical run, rather than a purely hands-off model evaluation.
Second-order effects
- Mathematicians evaluating AI on original problems gain another reason to distinguish a model's final answer from intermediate advances, an approach aligned with the unpublished-research test set.
- Other AI labs pursuing reasoning models face pressure to document research-task methodology—including prompting and human intervention—when presenting mathematical results as capability evidence.
Third-order effects
- If mathematics continues to serve as a leading reasoning benchmark, evaluation will shift toward reproducible human-model research workflows rather than isolated correct answers or failed headline goals.
- The combination of extended model runs and expert feedback points toward AI-assisted research systems in which progress is assessed as a collaborative process, not solely as autonomous theorem solving.
The trend: AI labs are using difficult mathematics to measure reasoning progress, with human-guided workflows becoming part of how those capabilities are demonstrated.