Researchers tested Google DeepMind's AlphaEvolve AI coding agent on 67 mathematical problems and found that it discovered improved solutions to ~20 of them
Really happy to share our new paper on using AlphaEvolve for mathematical exploration at scale, written with Javier Gómez-Serrano, Terence Tao, and @GoogleDeepMind's Bogdan Georgiev. We tested it on 67 problems and documented all our successes and failures. 🧵 [image]
Context & Ripple Effects
This is a more systematic test of AlphaEvolve after DeepMind introduced it as an evolutionary coding agent for algorithm design and optimization the AlphaEvolve launch. It extends a research line that included FunSearch's result on the cap set problem and specialized mathematical-reasoning systems.
The reported hit rate matters because it records both successes and failures across a defined problem set, rather than relying on a single standout result. That offers a more concrete basis for judging where human-AI mathematical collaboration may be useful.
First-order effects
- The participating researchers now have a documented evaluation of AlphaEvolve across 67 problems, including roughly 20 cases with improved solutions.
- For DeepMind, the result supplies evidence that its evolutionary-agent approach can contribute beyond the geometry and Olympiad-oriented capabilities previously associated with AlphaGeometry and AlphaProof DeepMind's math-specialist model rollout.
Second-order effects
- Mathematical AI teams face pressure to publish broader, failure-inclusive evaluations rather than foregrounding isolated discoveries, since comparative evidence is more useful to researchers choosing tools.
- Researchers may increasingly use coding agents to generate and refine candidate approaches, while reserving human effort for selecting problems, checking results, and assessing significance.
Third-order effects
- If replicated across other problem classes, mathematical AI could evolve from benchmark-focused solvers into a research workflow layer that expands the number of ideas experts can evaluate; the quality of verification and reporting will determine its practical value.
- The field may shift toward auditable human-AI research processes, where reproducible evaluation sets and documented failures matter as much as individual novel solutions.
The trend: AI mathematics is moving from narrowly scored problem-solving systems toward evaluated agents designed to augment exploratory research workflows.