/
Navigation
Chronicles
Browse all articles
Explore
Semantic exploration
Research
Entity momentum
Nexus
Correlations & relationships
Story Arc
Topic evolution
Drift Map
Semantic trajectory animation
Posts
Analysis & commentary
Pulse API
Tech news intelligence API
Browse
Entities
Companies, people, products, technologies
Domains
Browse by publication source
Handles
Browse by social media handle
Detection
Concept Search
Semantic similarity search
High Impact Stories
Top coverage by position
Sentiment Analysis
Positive/negative coverage
Anomaly Detection
Unusual coverage patterns
Analysis
Rivalry Report
Compare two entities head-to-head
Semantic Pivots
Narrative discontinuities
Crisis Response
Event recovery patterns
Connected
Search: /
Command: ⌘K
Embeddings: large
TEXXR

Chronicles

The story behind the story

days · browse · Enter similar · o open

Fields Medalist Timothy Gowers says most famous mathematics problems solved by LLMs so far have almost all been with counterexamples rather than proofs

Timothy Gowers /Gowers's Weblog:

Gowers's Weblog Timothy Gowers

Context & Ripple Effects

Recent coverage has ranged from DeepMind's FunSearch result on the cap set problem to mathematicians exploring how to use LLMs without losing direct mathematical understanding. Gowers's distinction matters because it separates finding a disproof from constructing a proof, two very different contributions to research.

The reported ChatGPT-assisted work on Crouzeix's conjecture has raised the visibility of LLMs in open-problem research; Gowers's assessment adds a narrower lens for judging what those successes demonstrate.

First-order effects

  • Mathematicians assessing LLM-assisted results will need to distinguish counterexample discovery from proof generation rather than treating both as the same kind of problem-solving advance.

Second-order effects

  • Researchers adopting LLMs in mathematical work are pushed toward workflows that preserve direct understanding, aligning with the concerns in work on incorporating LLMs into research without losing it.
  • Model developers face pressure to evaluate systems separately on generating valid proofs and on searching for counterexamples, since headline problem solves can mask that difference.

Third-order effects

  • If counterexamples continue to dominate prominent LLM mathematics results, the field may organize these tools primarily around conjecture testing and search while reserving proof construction for more tightly verified human-machine workflows.

The trend: LLM mathematics is moving from headline claims of solved problems toward capability-specific assessment of search, disproof, and proof generation.

Discussion

  • @thomasfbloom Thomas Bloom on x
    @littmath @wtgowers Yet this is exactly the sort of problem and solution strategy that AIs (and those prompting AIs) are most likely to favour.
  • @thomasfbloom Thomas Bloom on x
    @littmath @wtgowers Yes, excellent point. On a similar note, humans can be put off from “using standard technique/recent related work, push it a bit further, to make progress on a famous problem”, since there is a psychological element of “well if this worked others would have do…
  • @littmath Daniel Litt on x
    @wtgowers Good post! Let me propose another possible way to understand LLMs' relative advantage: I think the problems they are solving are ones where meaningful “partial progress” is relatively unlikely. Such problems are more expensive for humans to attempt...
  • @adi_baradwaj Adi on x
    Even as a non-mathematician, it's easy to notice that most of the recent AI-driven math results have been biased towards “examples” or “counterexamples”, rather than “theorems” Timothy Gowers has a great post about this, linked below
  • @wtgowers @wtgowers on x
    I've just written a blog post in which I discuss the question of whether current LLMs are better at some kinds of mathematics than others, with an emphasis on “current”, given how quickly the situation is developing. Link in next tweet.
  • @jmcnamara18 Jake McNamara on x
    Very interesting post, which I think also explains why we haven't yet seen the same level of results in physics as in math
  • @tcarmody Tim Carmody on bluesky
    This is such a basic point I wish I'd thought of it myself.  But I didn't [embedded post]