How a Claude model made a math breakthrough during its unsuccessful 54-hour attempt to solve the Riemann hypothesis after a user repeatedly encouraged it
Ben Cohen /Wall Street Journal:
Context & Ripple Effects
Anthropic had already disclosed that an unreleased Claude model failed on the Riemann hypothesis while making unexpected progress on a related problem in its earlier account of the Riemann effort. The Journal's account adds the operational detail of a prolonged, user-guided attempt, turning an outcome that could be framed simply as failure into a case study in how sustained model interaction can produce research-relevant intermediate results.
The result sits alongside recent claims of AI-assisted mathematical work, including an Anthropic mathematician's reported use of Fable 5 on the Jacobian conjecture, and a longer history of automated conjecture-generation systems such as the Ramanujan Machine.
First-order effects
- Anthropic gains a more concrete research-capability narrative for Claude: the model did not settle the target problem, but its extended attempt generated a reported mathematical advance.
- Mathematicians evaluating Claude now have to separate the reported related-problem result from the unresolved Riemann hypothesis, placing scrutiny on the novelty and verification of the intermediate work.
Second-order effects
- Competing model developers face a sharper benchmark than headline problem-solving: ChatGPT's reported role in work on Crouzeix's conjecture and Claude's related advance make externally checked mathematical contributions a differentiator.
- Researchers using frontier models are incentivized to treat long, iterative sessions as part of the research workflow rather than judge systems solely by whether they return a final solution to a posed problem.
Third-order effects
- If independently validated advances continue to emerge from failed target-task attempts, mathematical AI evaluation will shift toward systems that generate checkable conjectures, proofs, and subproblems—not binary claims of having solved famous open problems.
- The field would increasingly need norms that attribute human and model contributions separately, since extended prompting and expert verification are integral to the reported results.
The trend: Frontier AI is moving from automated conjecture generation toward human-guided, long-horizon mathematical research assistance measured by verifiable intermediate advances.