/
Navigation
Chronicles
Browse all articles
Explore
Semantic exploration
Research
Entity momentum
Nexus
Correlations & relationships
Story Arc
Topic evolution
Drift Map
Semantic trajectory animation
Posts
Analysis & commentary
Pulse API
Tech news intelligence API
Browse
Entities
Companies, people, products, technologies
Domains
Browse by publication source
Handles
Browse by social media handle
Detection
Concept Search
Semantic similarity search
High Impact Stories
Top coverage by position
Sentiment Analysis
Positive/negative coverage
Anomaly Detection
Unusual coverage patterns
Analysis
Rivalry Report
Compare two entities head-to-head
Semantic Pivots
Narrative discontinuities
Crisis Response
Event recovery patterns
Connected
Search: /
Command: ⌘K
Embeddings: large
TEXXR

Chronicles

The story behind the story

days · browse · Enter similar · o open

How a Claude model made a math breakthrough during its unsuccessful 54-hour attempt to solve the Riemann hypothesis after a user repeatedly encouraged it

Ben Cohen /Wall Street Journal:

Wall Street Journal Ben Cohen

Context & Ripple Effects

Anthropic had already disclosed that an unreleased Claude model failed on the Riemann hypothesis while making unexpected progress on a related problem in its earlier account of the Riemann effort. The Journal's account adds the operational detail of a prolonged, user-guided attempt, turning an outcome that could be framed simply as failure into a case study in how sustained model interaction can produce research-relevant intermediate results.

The result sits alongside recent claims of AI-assisted mathematical work, including an Anthropic mathematician's reported use of Fable 5 on the Jacobian conjecture, and a longer history of automated conjecture-generation systems such as the Ramanujan Machine.

First-order effects

  • Anthropic gains a more concrete research-capability narrative for Claude: the model did not settle the target problem, but its extended attempt generated a reported mathematical advance.
  • Mathematicians evaluating Claude now have to separate the reported related-problem result from the unresolved Riemann hypothesis, placing scrutiny on the novelty and verification of the intermediate work.

Second-order effects

  • Competing model developers face a sharper benchmark than headline problem-solving: ChatGPT's reported role in work on Crouzeix's conjecture and Claude's related advance make externally checked mathematical contributions a differentiator.
  • Researchers using frontier models are incentivized to treat long, iterative sessions as part of the research workflow rather than judge systems solely by whether they return a final solution to a posed problem.

Third-order effects

  • If independently validated advances continue to emerge from failed target-task attempts, mathematical AI evaluation will shift toward systems that generate checkable conjectures, proofs, and subproblems—not binary claims of having solved famous open problems.
  • The field would increasingly need norms that attribute human and model contributions separately, since extended prompting and expert verification are integral to the reported results.

The trend: Frontier AI is moving from automated conjecture generation toward human-guided, long-horizon mathematical research assistance measured by verifiable intermediate advances.