Fields Medalist Timothy Gowers says most famous mathematics problems solved by LLMs so far have almost all been with counterexamples rather than proofs
Timothy Gowers /Gowers's Weblog:
Context & Ripple Effects
Recent coverage has ranged from DeepMind's FunSearch result on the cap set problem to mathematicians exploring how to use LLMs without losing direct mathematical understanding. Gowers's distinction matters because it separates finding a disproof from constructing a proof, two very different contributions to research.
The reported ChatGPT-assisted work on Crouzeix's conjecture has raised the visibility of LLMs in open-problem research; Gowers's assessment adds a narrower lens for judging what those successes demonstrate.
First-order effects
- Mathematicians assessing LLM-assisted results will need to distinguish counterexample discovery from proof generation rather than treating both as the same kind of problem-solving advance.
Second-order effects
- Researchers adopting LLMs in mathematical work are pushed toward workflows that preserve direct understanding, aligning with the concerns in work on incorporating LLMs into research without losing it.
- Model developers face pressure to evaluate systems separately on generating valid proofs and on searching for counterexamples, since headline problem solves can mask that difference.
Third-order effects
- If counterexamples continue to dominate prominent LLM mathematics results, the field may organize these tools primarily around conjecture testing and search while reserving proof construction for more tightly verified human-machine workflows.
The trend: LLM mathematics is moving from headline claims of solved problems toward capability-specific assessment of search, disproof, and proof generation.