Fields Medalist Timothy Gowers says most famous mathematics problems solved by LLMs so far have almost all been with counterexamples rather than proofs
For the sake of anyone who might read this blog post in the distant future (a month from now, say), let me mention that I am writing …
Context & Ripple Effects
Recent coverage paired reports of a GPT-5.6-assisted solution to Crouzeix’s conjecture with mathematicians’ efforts to use LLMs without surrendering direct mathematical understanding. Gowers’s distinction matters because a counterexample can settle a conjecture without demonstrating the proof-generating capability implied by many “solved by AI” claims.
The record already includes DeepMind’s FunSearch result on the cap set problem and claims of olympiad-level reasoning performance. Gowers supplies a stricter lens for comparing such milestones: whether a system found a disproof or produced a proof.
First-order effects
- Reports of LLM-assisted mathematical breakthroughs, including the Crouzeix’s-conjecture claim, face more explicit scrutiny over whether the result is a counterexample or a proof.
- Mathematicians using LLMs gain a clearer criterion for documenting the model’s contribution: discovery of a falsifying instance versus construction and verification of a general argument.
Second-order effects
- LLM developers’ math-performance claims will be less comparable when they pool counterexample discovery with proof generation, pushing evaluators toward task-specific evidence.
- Research groups pursuing LLM-assisted mathematical workflows will need to preserve human access to the argument, rather than treating a correct outcome alone as evidence of mathematical understanding.
Third-order effects
- Mathematical AI evaluation is moving toward separating search and falsification from proof construction, a distinction that may determine which systems are trusted as research collaborators.
- If reported breakthroughs continue to rely disproportionately on counterexamples, the field’s progress narrative will center on automated conjecture testing before autonomous proof discovery.
The trend: AI mathematics is shifting from headline-level claims of “solving” problems toward finer evaluation of whether models disprove conjectures, find examples, or generate proofs.