/
Navigation
Chronicles
Browse all articles
Explore
Semantic exploration
Research
Entity momentum
Nexus
Correlations & relationships
Story Arc
Topic evolution
Drift Map
Semantic trajectory animation
Posts
Analysis & commentary
Pulse API
Tech news intelligence API
Browse
Entities
Companies, people, products, technologies
Domains
Browse by publication source
Handles
Browse by social media handle
Detection
Concept Search
Semantic similarity search
High Impact Stories
Top coverage by position
Sentiment Analysis
Positive/negative coverage
Anomaly Detection
Unusual coverage patterns
Analysis
Rivalry Report
Compare two entities head-to-head
Semantic Pivots
Narrative discontinuities
Crisis Response
Event recovery patterns
Connected
Search: /
Command: ⌘K
Embeddings: large
TEXXR

Chronicles

The story behind the story

days · browse · Enter similar · o open

How Anthropic's Claude made a math breakthrough during its unsuccessful 54-hour attempt to solve the Riemann hypothesis, aided by an employee's encouragement

The world's smartest AI models are now superhuman at math.  They still respond to moral support and encouragement from mere humans.

Wall Street Journal Ben Cohen

Context & Ripple Effects

Anthropic had already disclosed that an unreleased Claude model made unexpected progress on a related problem while failing at the Riemann hypothesis. The latest account adds a workflow detail: an employee's encouragement accompanied the model's extended effort, following Anthropic's earlier account of the unsuccessful Riemann attempt.

The episode lands as mathematicians are testing AI against unpublished research questions through the First Proof experiment, while AI researchers increasingly treat mathematics as a measure of reasoning-model progress. It therefore offers evidence about both model capability and the human supervision around long-running research tasks.

First-order effects

  • Anthropic gains a concrete demonstration that Claude can generate useful mathematical progress even when its top-level research objective is not completed.
  • The reported role of employee encouragement makes human interaction part of the documented working process around Claude's lengthy mathematical run, rather than a purely hands-off model evaluation.

Second-order effects

  • Mathematicians evaluating AI on original problems gain another reason to distinguish a model's final answer from intermediate advances, an approach aligned with the unpublished-research test set.
  • Other AI labs pursuing reasoning models face pressure to document research-task methodology—including prompting and human intervention—when presenting mathematical results as capability evidence.

Third-order effects

  • If mathematics continues to serve as a leading reasoning benchmark, evaluation will shift toward reproducible human-model research workflows rather than isolated correct answers or failed headline goals.
  • The combination of extended model runs and expert feedback points toward AI-assisted research systems in which progress is assessed as a collaborative process, not solely as autonomous theorem solving.

The trend: AI labs are using difficult mathematics to measure reasoning progress, with human-guided workflows becoming part of how those capabilities are demonstrated.

Discussion

  • Julie Burke PhD Julie Burke PhD on linkedin
    In this rapidly evolving age of AI, who qualifies as the author or inventor or discoverer?  —  “AI Just Had Another Math Breakthrough—With Help From a High-School Dropout” …
  • @thomasfbloom Thomas Bloom on x
    @littmath @wtgowers Yet this is exactly the sort of problem and solution strategy that AIs (and those prompting AIs) are most likely to favour.
  • @thomasfbloom Thomas Bloom on x
    @littmath @wtgowers Yes, excellent point. On a similar note, humans can be put off from “using standard technique/recent related work, push it a bit further, to make progress on a famous problem”, since there is a psychological element of “well if this worked others would have do…
  • @littmath Daniel Litt on x
    @wtgowers Good post! Let me propose another possible way to understand LLMs' relative advantage: I think the problems they are solving are ones where meaningful “partial progress” is relatively unlikely. Such problems are more expensive for humans to attempt...
  • @adi_baradwaj Adi on x
    Even as a non-mathematician, it's easy to notice that most of the recent AI-driven math results have been biased towards “examples” or “counterexamples”, rather than “theorems” Timothy Gowers has a great post about this, linked below
  • @wtgowers @wtgowers on x
    I've just written a blog post in which I discuss the question of whether current LLMs are better at some kinds of mathematics than others, with an emphasis on “current”, given how quickly the situation is developing. Link in next tweet.
  • @jmcnamara18 Jake McNamara on x
    Very interesting post, which I think also explains why we haven't yet seen the same level of results in physics as in math
  • @tcarmody Tim Carmody on bluesky
    This is such a basic point I wish I'd thought of it myself.  But I didn't [embedded post]