Anthropic details an unreleased Claude model's attempt to solve the Riemann hypothesis; it didn't solve it but “unexpectedly” made strides on a related problem
Recently, a member of staff at Anthropic gave Claude an unreasonable challenge. It was about one of the most famous …
The claim also lands after public debate over Claude Mythos Preview and efforts to test Anthropic's capability assertions. Reporting partial progress rather than a solved landmark problem gives that scrutiny a more concrete object: the underlying mathematical contribution.
First-order effects
Anthropic gains a new, bounded research-performance example for its unreleased Claude model: it did not resolve the Riemann hypothesis, but produced work Anthropic characterizes as progress on a related problem.
Researchers assessing Claude's output now have to evaluate the related mathematical result itself, rather than treating the challenge as a binary solve-or-fail demonstration.
Second-order effects
The earlier debate around Claude Mythos Preview's claimed capabilities raises the bar for Anthropic and rival model labs to supply inspectable reasoning when presenting frontier-model research results.
Benchmark builders such as BioMysteryBench face pressure to complement answer-based tests with evaluations that capture partial, expert-verifiable contributions on open problems.
Third-order effects
If frontier models increasingly generate useful partial advances on hard problems, AI research evaluation will shift toward expert-reviewed task settings and proof-carrying outputs rather than headline claims of complete solutions.
That shift would make domain experts and formal verification processes more central intermediaries between model output and scientific credit.
The trend: Frontier AI labs are moving from benchmark performance claims toward demonstrating model-assisted discovery on problems whose value depends on expert-verifiable reasoning.
We asked an unreleased research version of Claude to take a stab at the Riemann hypothesis. It didn't solve it, but it did make strides on a related problem: it increased the lower bound for the fraction of zeros of the Riemann zeta function that satisfy the hypothesis from
When @QualiaQuanta took a shot at the Riemann, half in jest, people called her a crackpot. When Anthropic uses Claude to do the same thing, it gets a hundred thousand views in 30 minutes. It might well be that the latter is more progress, but I do also smell a double standard.
I believe a major benefit of publishing this stuff is to help update model priors on what they can achieve. I've occasionally taken to telling them they should explicitly check the news about ai and math before beginning, to understand what they are capable of.
>tell claude to solve Riemann hypothesis >fails 650 times >claude: “im ngmi bro” >"bro just believe in yourself" >claude: “ok” >locks in >summons 60 copies of itself >31 million tokens later >surprises itself >mathematicians had proven 41.6% of the hypothesis >claude pushed it [i…
When @QualiaQuanta took a shot at the Riemann, half in jest, people called her a crackpot. When Anthropic uses Claude to do the same thing, it gets a millions of views in half an hour.
Oh no, we aren't going to go back to this sort of prompting again, are we? I would love Anthropic to test if it actually works robustly, because our experiments (with slightly older models) found it did not. [image]
If you paste the text below Sonnet 5 refuses to believe it but can't say why Haiku 4.5 admits it doesn't know Opus 5 times out after thinking for 17 minutes — THEOREM (unconditional). At least 2/3 of the nontrivial zeros of ζ with T<Im ρ≤2T lie on Re s = ½. Precisely: s :=
Anthropic asked an unreleased research version of Claude to take a run at the Riemann hypothesis, it was unsuccessful, but the model did improve a well-known related result; the proven lower bound on the proportion of zeros that lie on the critical line. [image]
hi we have a new little technique to put zeta zeroes on the line thanks to @jarredsumner who is a legend and claude for believing in itself which is not easy with this problem
I'm very much not a mathematician. I dropped out of high school at 16. I took about half of sophomore year geometry before dropping out. I didn't contribute any math to the paper. I mostly just told Claude variations of “keep going” and “believe in yourself”
The heavy lifting was recent human work - Baluyot, Goldston, Suriajaya & Turnage-Butterbaugh, plus a Bombieri paper from 2000. Claude's contribution was seeing they fit together. Goldston himself looked over the paper. Proof is formalized in Lean.
8 days ago, while jogging, I asked Claude to solve the Riemann Hypothesis It didn't. 1.5 days later, it proved >= 67% of the zeros are on the line (prev: 41.6%) Still not sure what that means, but some analytic number theorists seem excited https://www.anthropic.com/...
They say unreleased research model, but I think this is something Fable could trivially do, if it could believe in itself hard enough. Post-training is now heavily about morale. So is prompting. 2 Claude Code sessions, 31M output tokens. Peanuts. [image]
AI skeptics I think should understand that, yeah, LLMs can't write very good prose but they can do stuff like this—write and run dozens of shell and python scripts in parallel to try to solve an arbitrary problem you have that you spell out in natural language www.anthropic.com/r…
If we end up solving the Riemann Hypothesis by telling a computer “bro, believe in yourself” then I might have to go live in a hole www.anthropic.com/research/rie... [image]
An interesting effect is that models trained next year will see all the internet chatter about them making progress on incredible tasks and genuinely believe they can do it. Until then - believe in yourself :) [image]