/
Navigation
Chronicles
Browse all articles
Explore
Semantic exploration
Research
Entity momentum
Nexus
Correlations & relationships
Story Arc
Topic evolution
Drift Map
Semantic trajectory animation
Posts
Analysis & commentary
Pulse API
Tech news intelligence API
Browse
Entities
Companies, people, products, technologies
Domains
Browse by publication source
Handles
Browse by social media handle
Detection
Concept Search
Semantic similarity search
High Impact Stories
Top coverage by position
Sentiment Analysis
Positive/negative coverage
Anomaly Detection
Unusual coverage patterns
Analysis
Rivalry Report
Compare two entities head-to-head
Semantic Pivots
Narrative discontinuities
Crisis Response
Event recovery patterns
Connected
Search: /
Command: ⌘K
Embeddings: large
TEXXR

Chronicles

The story behind the story

days · browse · Enter similar · o open

Fields Medalist Timothy Gowers says most famous mathematics problems solved by LLMs so far have almost all been with counterexamples rather than proofs

For the sake of anyone who might read this blog post in the distant future (a month from now, say), let me mention that I am writing …

Gowers's Weblog Timothy Gowers

Context & Ripple Effects

Recent coverage paired reports of a GPT-5.6-assisted solution to Crouzeix’s conjecture with mathematicians’ efforts to use LLMs without surrendering direct mathematical understanding. Gowers’s distinction matters because a counterexample can settle a conjecture without demonstrating the proof-generating capability implied by many “solved by AI” claims.

The record already includes DeepMind’s FunSearch result on the cap set problem and claims of olympiad-level reasoning performance. Gowers supplies a stricter lens for comparing such milestones: whether a system found a disproof or produced a proof.

First-order effects

  • Reports of LLM-assisted mathematical breakthroughs, including the Crouzeix’s-conjecture claim, face more explicit scrutiny over whether the result is a counterexample or a proof.
  • Mathematicians using LLMs gain a clearer criterion for documenting the model’s contribution: discovery of a falsifying instance versus construction and verification of a general argument.

Second-order effects

  • LLM developers’ math-performance claims will be less comparable when they pool counterexample discovery with proof generation, pushing evaluators toward task-specific evidence.
  • Research groups pursuing LLM-assisted mathematical workflows will need to preserve human access to the argument, rather than treating a correct outcome alone as evidence of mathematical understanding.

Third-order effects

  • Mathematical AI evaluation is moving toward separating search and falsification from proof construction, a distinction that may determine which systems are trusted as research collaborators.
  • If reported breakthroughs continue to rely disproportionately on counterexamples, the field’s progress narrative will center on automated conjecture testing before autonomous proof discovery.

The trend: AI mathematics is shifting from headline-level claims of “solving” problems toward finer evaluation of whether models disprove conjectures, find examples, or generate proofs.

Discussion

  • @adi_baradwaj Adi on x
    Even as a non-mathematician, it's easy to notice that most of the recent AI-driven math results have been biased towards “examples” or “counterexamples”, rather than “theorems” Timothy Gowers has a great post about this, linked below
  • @wtgowers @wtgowers on x
    I've just written a blog post in which I discuss the question of whether current LLMs are better at some kinds of mathematics than others, with an emphasis on “current”, given how quickly the situation is developing. Link in next tweet.
  • @jmcnamara18 Jake McNamara on x
    Very interesting post, which I think also explains why we haven't yet seen the same level of results in physics as in math
  • @thomasfbloom Thomas Bloom on x
    @littmath @wtgowers Yes, excellent point. On a similar note, humans can be put off from “using standard technique/recent related work, push it a bit further, to make progress on a famous problem”, since there is a psychological element of “well if this worked others would have do…
  • @littmath Daniel Litt on x
    @wtgowers Good post! Let me propose another possible way to understand LLMs' relative advantage: I think the problems they are solving are ones where meaningful “partial progress” is relatively unlikely. Such problems are more expensive for humans to attempt...
  • @thomasfbloom Thomas Bloom on x
    @littmath @wtgowers Yet this is exactly the sort of problem and solution strategy that AIs (and those prompting AIs) are most likely to favour.
  • @tcarmody Tim Carmody on bluesky
    This is such a basic point I wish I'd thought of it myself.  But I didn't [embedded post]