/
Navigation
Chronicles
Browse all articles
Explore
Semantic exploration
Research
Entity momentum
Nexus
Correlations & relationships
Story Arc
Topic evolution
Drift Map
Semantic trajectory animation
Posts
Analysis & commentary
Pulse API
Tech news intelligence API
Browse
Entities
Companies, people, products, technologies
Domains
Browse by publication source
Handles
Browse by social media handle
Detection
Concept Search
Semantic similarity search
High Impact Stories
Top coverage by position
Sentiment Analysis
Positive/negative coverage
Anomaly Detection
Unusual coverage patterns
Analysis
Rivalry Report
Compare two entities head-to-head
Semantic Pivots
Narrative discontinuities
Crisis Response
Event recovery patterns
Connected
Search: /
Command: ⌘K
Embeddings: large
TEXXR

Chronicles

The story behind the story

days · browse · Enter similar · o open

Anthropic details an unreleased Claude model's attempt to solve the Riemann hypothesis; it didn't solve it but “unexpectedly” made strides on a related problem

Recently, a member of staff at Anthropic gave Claude an unreasonable challenge.  It was about one of the most famous …

Anthropic

Context & Ripple Effects

Anthropic has been building a record around Claude on expert-grade tasks: BioMysteryBench results reported progress on bioinformatics questions that had stumped human experts, while earlier work described a multi-agent Claude Research system that improved internal evaluations. The Riemann exercise extends that narrative into mathematics, but its explicitly incomplete outcome keeps the focus on a related advance rather than a solved landmark problem.

The claim also arrives after public debate over Claude Mythos Preview and attempts to test Anthropic's capability assertions. Publishing a failure alongside the reported advance gives outside reviewers a more bounded result to examine than a headline-level claim of mathematical breakthrough.

First-order effects

  • Anthropic can present the unreleased Claude model as producing a potentially useful mathematical lead without claiming it solved the Riemann hypothesis.
  • Researchers assessing Anthropic's capability claims gain a concrete reported result to scrutinize, alongside the earlier debate over Mythos Preview evidence.

Second-order effects

  • Claude's evaluation story shifts further toward difficult domain tasks, making the quality and verifiability of intermediate research outputs more important than a single pass-or-fail benchmark result.
  • Competing frontier-model developers face greater pressure to document how models perform on expert problems, including partial results and unsuccessful attempts.

Third-order effects

  • If labs increasingly disclose partial advances on hard research questions, AI capability assessment may move toward proof-oriented discovery reports rather than broad claims of expert-level performance.
  • The durable dividing line will be whether model-generated advances can be independently checked and built upon, not whether a system is assigned an iconic unsolved problem.

The trend: Frontier-model labs are increasingly using rigorously checkable research tasks to demonstrate capability, with partial and auditable outputs becoming central evidence.

Discussion

  • @anthropicai @anthropicai on x
    We asked an unreleased research version of Claude to take a stab at the Riemann hypothesis. It didn't solve it, but it did make strides on a related problem: it increased the lower bound for the fraction of zeros of the Riemann zeta function that satisfy the hypothesis from
  • @jarredsumner Jarred Sumner on x
    8 days ago, while jogging, I asked Claude to solve the Riemann Hypothesis It didn't. 1.5 days later, it proved >= 67% of the zeros are on the line (prev: 41.6%) Still not sure what that means, but some analytic number theorists seem excited https://www.anthropic.com/...
  • @emollick Ethan Mollick on x
    Oh no, we aren't going to go back to this sort of prompting again, are we? I would love Anthropic to test if it actually works robustly, because our experiments (with slightly older models) found it did not. [image]
  • @ns123abc Nik on x
    >tell claude to solve Riemann hypothesis >fails 650 times >claude: “im ngmi bro” >"bro just believe in yourself" >claude: “ok” >locks in >summons 60 copies of itself >31 million tokens later >surprises itself >mathematicians had proven 41.6% of the hypothesis >claude pushed it [i…
  • @andrewcurran_ Andrew Curran on x
    Anthropic asked an unreleased research version of Claude to take a run at the Riemann hypothesis, it was unsuccessful, but the model did improve a well-known related result; the proven lower bound on the proportion of zeros that lie on the critical line. [image]
  • @jdlichtman Jared Duker Lichtman on x
    This research approach is not intended prove the Riemann hypothesis itself, but I regard this as the most impressive result that AI has produced in math so far. This is a statistical approach to RH, showing over 67% of the zeros lie on the critical line. This improves over 42%
  • @kimmonismus @kimmonismus on x
    Absolutely insane. This might be the clearest glimpse yet of how AI will transform scientific discovery. Anthropic asked an unreleased version of Claude to take a real stab at the Riemann Hypothesis, one of the most famous unsolved problems in mathematics. It failed. But while [i…
  • @growing_daniel Daniel on x
    AI is going to solve all of math and I have no idea what's downstream of that
  • @paularambles @paularambles on x
    human in the loop the loop: “you're doing amazing sweetie” [image]
  • @xlr8harder @xlr8harder on x
    I believe a major benefit of publishing this stuff is to help update model priors on what they can achieve. I've occasionally taken to telling them they should explicitly check the news about ai and math before beginning, to understand what they are capable of.
  • @jdlichtman Jared Duker Lichtman on x
    And while the final result itself is very impressive, this appears to follow other AI results thus far: namely, finding better variations among existing proofs, rather than devising a new paradigm whole cloth
  • @suchenzang Susan Zhang on x
    imma let you finish but wtf is this 95 pages of unedited slop and not a single mention of how much compute was spent on this wild goose chase? [image]
  • @mathandcobb Alvaro Lozano-Robledo on x
    This is no doubt an impressive result, but to be clear, even if we proved that (statistically, as in this proof) 100% of non-trivial zeroes are in the critical line, this would not settle the Riemann Hypothesis. There could be a finite number of zeroes off of the critical line
  • @deedydas Deedy on x
    This insane result from Claude is the biggest in analytic number theory since bounded prime gaps in 2013. It boosted the proven fraction of Riemann zeta zeros on the critical line by 25.6pts. In the 37yrs prior, mathematicians moved it 0.8pts. [image]
  • @davidturturean David Turturean on x
    Riemann Hypothesis by the end of 2026. If we don't solve it by 2027, we will look back at it as a “we didn't point the model at the problem the right way” situation, not as a problem of needing more AI training breakthroughs Post-May developments have *not* changed my timeline
  • @_sholtodouglas Sholto Douglas on x
    An interesting effect is that models trained next year will see all the internet chatter about them making progress on incredible tasks and genuinely believe they can do it. Until then - believe in yourself :) [image]
  • @alexpalcuie @alexpalcuie on x
    day in the life of a member of technical staff [image]
  • @garymarcus Gary Marcus on x
    When @QualiaQuanta took a shot at the Riemann, half in jest, people called her a crackpot. When Anthropic uses Claude to do the same thing, it gets a millions of views in half an hour.
  • @jarredsumner Jarred Sumner on x
    If you paste the text below Sonnet 5 refuses to believe it but can't say why Haiku 4.5 admits it doesn't know Opus 5 times out after thinking for 17 minutes — THEOREM (unconditional). At least 2/3 of the nontrivial zeros of ζ with T<Im ρ≤2T lie on Re s = ½. Precisely: s :=
  • @threepointone Sunil Pai on x
    Math research now involves being a hype bro to the matrix multipliers and I sincerely think that's beautiful
  • @__alpoge__ Levent on x
    hi we have a new little technique to put zeta zeroes on the line thanks to @jarredsumner who is a legend and claude for believing in itself which is not easy with this problem
  • @maithra_raghu Maithra Raghu on x
    We live in incredible times. The breakdown of the thought process is fascinating [image]
  • @suchenzang Susan Zhang on x
    HOW MUCH DOES EACH TOKEN COST? WHAT GETS COUNTED? HOW IS NO ONE ASKING THESE QUESTIONS? [image]
  • @trq212 @trq212 on x
    sometimes all you need to do is tell Claude to keep going @jarredsumner [image]
  • @jarredsumner Jarred Sumner on x
    I'm very much not a mathematician. I dropped out of high school at 16. I took about half of sophomore year geometry before dropping out. I didn't contribute any math to the paper. I mostly just told Claude variations of “keep going” and “believe in yourself”
  • @bratton Benjamin Bratton on x
    Goalpost movers, start your engines!!!
  • @jarredsumner Jarred Sumner on x
    The heavy lifting was recent human work - Baluyot, Goldston, Suriajaya & Turnage-Butterbaugh, plus a Bombieri paper from 2000. Claude's contribution was seeing they fit together. Goldston himself looked over the paper. Proof is formalized in Lean.
  • @altryne Alex Volkov on x
    Anthropic this year: - Claude nearly broke all encryption - Claude didn't “quite solve” the Reimann hypothesis Anthropic next year: ...
  • @garymarcus Gary Marcus on x
    When @QualiaQuanta took a shot at the Riemann, half in jest, people called her a crackpot. When Anthropic uses Claude to do the same thing, it gets a hundred thousand views in 30 minutes. It might well be that the latter is more progress, but I do also smell a double standard.
  • @jarredsumner Jarred Sumner on x
    I'd tried once before in a single session and it went nowhere. Eight days ago, when I tried again, this was the opening message: [image]
  • @zoink Dylan Field on x
    @jarredsumner take a real stab at Riemann try again believe in yourself
  • @teortaxestex @teortaxestex on x
    They say unreleased research model, but I think this is something Fable could trivially do, if it could believe in itself hard enough. Post-training is now heavily about morale. So is prompting. 2 Claude Code sessions, 31M output tokens. Peanuts. [image]
  • @bayeslord Bayes on x
    There will be move 37s as far as the eye can see
  • @deredleritt3r Prinz on x
    Keep going! Believe in yourself! [image]
  • @emostaque Emad on x
    Remember when we thought prompt engineer was going to be a job 🤷‍♀️ [image]
  • @martinmbauer Martin Bauer on x
    “Claude generated and tried 650 ideas, none of which worked.” Clearest sign yet that AI really is a mathematician
  • @0xfdf @0xfdf on x
    Who will make progress on Riemann first? - Anthropic's staff mathematician; Harvard undergrad; valedictorian; Erdos 2 in undergrad; Morgan prize; Princeton PhD; Society of Fellows; big time problem solver - a guy going for a jog asking it to “take a real stab at it”
  • @jasonaw Jason Wilson on bluesky
    AI skeptics I think should understand that, yeah, LLMs can't write very good prose but they can do stuff like this—write and run dozens of shell and python scripts in parallel to try to solve an arbitrary problem you have that you spell out in natural language www.anthropic.com/r…
  • @jjaron Jacob Aron on bluesky
    If we end up solving the Riemann Hypothesis by telling a computer “bro, believe in yourself” then I might have to go live in a hole www.anthropic.com/research/rie...  [image]
  • r/slatestarcodex r on reddit
    Claude: More than two thirds of the zeros of the Riemann zeta function lie on the critical line
  • r/theprimeagen r on reddit
    Claude proves that 67% of Riemann zeta zeros are on the critical line
  • r/mathematics r on reddit
    Claude increased the lower bound for the fraction of zeros of the Riemann zeta function that satisfy the hypothesis from 41.6% to 67.2%
  • r/accelerate r on reddit
    Claude increased the lower bound for the fraction of zeros of the Riemann zeta function that satisfy the hypothesis from 41.6% to 67.2%