Anthropic details an unreleased Claude model's attempt to solve the Riemann hypothesis; it didn't solve it but “unexpectedly” made strides on a related problem
Recently, a member of staff at Anthropic gave Claude an unreasonable challenge. It was about one of the most famous …
Anthropic
Context & Ripple Effects
Anthropic has been building a record around Claude on expert-grade tasks: BioMysteryBench results reported progress on bioinformatics questions that had stumped human experts, while earlier work described a multi-agent Claude Research system that improved internal evaluations. The Riemann exercise extends that narrative into mathematics, but its explicitly incomplete outcome keeps the focus on a related advance rather than a solved landmark problem.
The claim also arrives after public debate over Claude Mythos Preview and attempts to test Anthropic's capability assertions. Publishing a failure alongside the reported advance gives outside reviewers a more bounded result to examine than a headline-level claim of mathematical breakthrough.
First-order effects
- Anthropic can present the unreleased Claude model as producing a potentially useful mathematical lead without claiming it solved the Riemann hypothesis.
- Researchers assessing Anthropic's capability claims gain a concrete reported result to scrutinize, alongside the earlier debate over Mythos Preview evidence.
Second-order effects
- Claude's evaluation story shifts further toward difficult domain tasks, making the quality and verifiability of intermediate research outputs more important than a single pass-or-fail benchmark result.
- Competing frontier-model developers face greater pressure to document how models perform on expert problems, including partial results and unsuccessful attempts.
Third-order effects
- If labs increasingly disclose partial advances on hard research questions, AI capability assessment may move toward proof-oriented discovery reports rather than broad claims of expert-level performance.
- The durable dividing line will be whether model-generated advances can be independently checked and built upon, not whether a system is assigned an iconic unsolved problem.
The trend: Frontier-model labs are increasingly using rigorously checkable research tasks to demonstrate capability, with partial and auditable outputs becoming central evidence.
Related: Proof-carrying discovery · Anthropic · Claude · Claude's BioMysteryBench results · The debate over Claude Mythos Preview · Anthropic's multi-agent Claude Research system
Related Coverage
- Unreleased Claude model makes breakthrough on century-old Riemann Hypothesis math problem Neowin · Paul Hill
- MORE THAN TWO THIRDS OF THE ZEROS OF THE RIEMANN ZETA FUNCTION LIE ON THE CRITICAL LINE Anthropic
- Zeta23 — a Lean 4 formalization of “More than two thirds of the zeros of the Riemann zeta function lie on the critical line” GitHub
- Claude AI Advances Riemann Zeta Problem Bound to 67.2% Blockchain.News · Terrill Dicki
- Learning more about Claude's mathematical capabilities Hacker News
- More than two thirds of the zeros of the Riemann zeta function lie on the critical line Lobsters
- Tenacious AI agents expose dark side of machine autonomy Axios · Zachary Basu
- An unreleased Anthropic model made progress on one of math's biggest unsolved problems TechCrunch · Russell Brandom
- Meta's Big Open Source Comeback Superintelligence
Discussion
-
@anthropicai
@anthropicai
on x
We asked an unreleased research version of Claude to take a stab at the Riemann hypothesis. It didn't solve it, but it did make strides on a related problem: it increased the lower bound for the fraction of zeros of the Riemann zeta function that satisfy the hypothesis from
-
@jarredsumner
Jarred Sumner
on x
8 days ago, while jogging, I asked Claude to solve the Riemann Hypothesis It didn't. 1.5 days later, it proved >= 67% of the zeros are on the line (prev: 41.6%) Still not sure what that means, but some analytic number theorists seem excited https://www.anthropic.com/...
-
@emollick
Ethan Mollick
on x
Oh no, we aren't going to go back to this sort of prompting again, are we? I would love Anthropic to test if it actually works robustly, because our experiments (with slightly older models) found it did not. [image]
-
@ns123abc
Nik
on x
>tell claude to solve Riemann hypothesis >fails 650 times >claude: “im ngmi bro” >"bro just believe in yourself" >claude: “ok” >locks in >summons 60 copies of itself >31 million tokens later >surprises itself >mathematicians had proven 41.6% of the hypothesis >claude pushed it [i…
-
@andrewcurran_
Andrew Curran
on x
Anthropic asked an unreleased research version of Claude to take a run at the Riemann hypothesis, it was unsuccessful, but the model did improve a well-known related result; the proven lower bound on the proportion of zeros that lie on the critical line. [image]
-
@jdlichtman
Jared Duker Lichtman
on x
This research approach is not intended prove the Riemann hypothesis itself, but I regard this as the most impressive result that AI has produced in math so far. This is a statistical approach to RH, showing over 67% of the zeros lie on the critical line. This improves over 42%
-
@kimmonismus
@kimmonismus
on x
Absolutely insane. This might be the clearest glimpse yet of how AI will transform scientific discovery. Anthropic asked an unreleased version of Claude to take a real stab at the Riemann Hypothesis, one of the most famous unsolved problems in mathematics. It failed. But while [i…
-
@growing_daniel
Daniel
on x
AI is going to solve all of math and I have no idea what's downstream of that
-
@paularambles
@paularambles
on x
human in the loop the loop: “you're doing amazing sweetie” [image]
-
@xlr8harder
@xlr8harder
on x
I believe a major benefit of publishing this stuff is to help update model priors on what they can achieve. I've occasionally taken to telling them they should explicitly check the news about ai and math before beginning, to understand what they are capable of.
-
@jdlichtman
Jared Duker Lichtman
on x
And while the final result itself is very impressive, this appears to follow other AI results thus far: namely, finding better variations among existing proofs, rather than devising a new paradigm whole cloth
-
@suchenzang
Susan Zhang
on x
imma let you finish but wtf is this 95 pages of unedited slop and not a single mention of how much compute was spent on this wild goose chase? [image]
-
@mathandcobb
Alvaro Lozano-Robledo
on x
This is no doubt an impressive result, but to be clear, even if we proved that (statistically, as in this proof) 100% of non-trivial zeroes are in the critical line, this would not settle the Riemann Hypothesis. There could be a finite number of zeroes off of the critical line
-
@deedydas
Deedy
on x
This insane result from Claude is the biggest in analytic number theory since bounded prime gaps in 2013. It boosted the proven fraction of Riemann zeta zeros on the critical line by 25.6pts. In the 37yrs prior, mathematicians moved it 0.8pts. [image]
-
@davidturturean
David Turturean
on x
Riemann Hypothesis by the end of 2026. If we don't solve it by 2027, we will look back at it as a “we didn't point the model at the problem the right way” situation, not as a problem of needing more AI training breakthroughs Post-May developments have *not* changed my timeline
-
@_sholtodouglas
Sholto Douglas
on x
An interesting effect is that models trained next year will see all the internet chatter about them making progress on incredible tasks and genuinely believe they can do it. Until then - believe in yourself :) [image]
-
@alexpalcuie
@alexpalcuie
on x
day in the life of a member of technical staff [image]
-
@garymarcus
Gary Marcus
on x
When @QualiaQuanta took a shot at the Riemann, half in jest, people called her a crackpot. When Anthropic uses Claude to do the same thing, it gets a millions of views in half an hour.
-
@jarredsumner
Jarred Sumner
on x
If you paste the text below Sonnet 5 refuses to believe it but can't say why Haiku 4.5 admits it doesn't know Opus 5 times out after thinking for 17 minutes — THEOREM (unconditional). At least 2/3 of the nontrivial zeros of ζ with T<Im ρ≤2T lie on Re s = ½. Precisely: s :=
-
@threepointone
Sunil Pai
on x
Math research now involves being a hype bro to the matrix multipliers and I sincerely think that's beautiful
-
@__alpoge__
Levent
on x
hi we have a new little technique to put zeta zeroes on the line thanks to @jarredsumner who is a legend and claude for believing in itself which is not easy with this problem
-
@maithra_raghu
Maithra Raghu
on x
We live in incredible times. The breakdown of the thought process is fascinating [image]
-
@suchenzang
Susan Zhang
on x
HOW MUCH DOES EACH TOKEN COST? WHAT GETS COUNTED? HOW IS NO ONE ASKING THESE QUESTIONS? [image]
-
@trq212
@trq212
on x
sometimes all you need to do is tell Claude to keep going @jarredsumner [image]
-
@jarredsumner
Jarred Sumner
on x
I'm very much not a mathematician. I dropped out of high school at 16. I took about half of sophomore year geometry before dropping out. I didn't contribute any math to the paper. I mostly just told Claude variations of “keep going” and “believe in yourself”
-
@bratton
Benjamin Bratton
on x
Goalpost movers, start your engines!!!
-
@jarredsumner
Jarred Sumner
on x
The heavy lifting was recent human work - Baluyot, Goldston, Suriajaya & Turnage-Butterbaugh, plus a Bombieri paper from 2000. Claude's contribution was seeing they fit together. Goldston himself looked over the paper. Proof is formalized in Lean.
-
@altryne
Alex Volkov
on x
Anthropic this year: - Claude nearly broke all encryption - Claude didn't “quite solve” the Reimann hypothesis Anthropic next year: ...
-
@garymarcus
Gary Marcus
on x
When @QualiaQuanta took a shot at the Riemann, half in jest, people called her a crackpot. When Anthropic uses Claude to do the same thing, it gets a hundred thousand views in 30 minutes. It might well be that the latter is more progress, but I do also smell a double standard.
-
@jarredsumner
Jarred Sumner
on x
I'd tried once before in a single session and it went nowhere. Eight days ago, when I tried again, this was the opening message: [image]
-
@zoink
Dylan Field
on x
@jarredsumner take a real stab at Riemann try again believe in yourself
-
@teortaxestex
@teortaxestex
on x
They say unreleased research model, but I think this is something Fable could trivially do, if it could believe in itself hard enough. Post-training is now heavily about morale. So is prompting. 2 Claude Code sessions, 31M output tokens. Peanuts. [image]
-
@bayeslord
Bayes
on x
There will be move 37s as far as the eye can see
-
@deredleritt3r
Prinz
on x
Keep going! Believe in yourself! [image]
-
@emostaque
Emad
on x
Remember when we thought prompt engineer was going to be a job 🤷♀️ [image]
-
@martinmbauer
Martin Bauer
on x
“Claude generated and tried 650 ideas, none of which worked.” Clearest sign yet that AI really is a mathematician
-
@0xfdf
@0xfdf
on x
Who will make progress on Riemann first? - Anthropic's staff mathematician; Harvard undergrad; valedictorian; Erdos 2 in undergrad; Morgan prize; Princeton PhD; Society of Fellows; big time problem solver - a guy going for a jog asking it to “take a real stab at it”
-
@jasonaw
Jason Wilson
on bluesky
AI skeptics I think should understand that, yeah, LLMs can't write very good prose but they can do stuff like this—write and run dozens of shell and python scripts in parallel to try to solve an arbitrary problem you have that you spell out in natural language www.anthropic.com/r…
-
@jjaron
Jacob Aron
on bluesky
If we end up solving the Riemann Hypothesis by telling a computer “bro, believe in yourself” then I might have to go live in a hole www.anthropic.com/research/rie... [image]
-
r/slatestarcodex
r
on reddit
Claude: More than two thirds of the zeros of the Riemann zeta function lie on the critical line
-
r/theprimeagen
r
on reddit
Claude proves that 67% of Riemann zeta zeros are on the critical line
-
r/mathematics
r
on reddit
Claude increased the lower bound for the fraction of zeros of the Riemann zeta function that satisfy the hypothesis from 41.6% to 67.2%
-
r/accelerate
r
on reddit
Claude increased the lower bound for the fraction of zeros of the Riemann zeta function that satisfy the hypothesis from 41.6% to 67.2%