/
Navigation
Chronicles
Browse all articles
Explore
Semantic exploration
Research
Entity momentum
Nexus
Correlations & relationships
Story Arc
Topic evolution
Drift Map
Semantic trajectory animation
Posts
Analysis & commentary
Pulse API
Tech news intelligence API
Browse
Entities
Companies, people, products, technologies
Domains
Browse by publication source
Handles
Browse by social media handle
Detection
Concept Search
Semantic similarity search
High Impact Stories
Top coverage by position
Sentiment Analysis
Positive/negative coverage
Anomaly Detection
Unusual coverage patterns
Analysis
Rivalry Report
Compare two entities head-to-head
Semantic Pivots
Narrative discontinuities
Crisis Response
Event recovery patterns
Connected
Search: /
Command: ⌘K
Embeddings: large
TEXXR

Chronicles

The story behind the story

days · browse · Enter similar · o open

Researchers tested Google DeepMind's AlphaEvolve AI coding agent on 67 mathematical problems and found that it discovered improved solutions to ~20 of them

Really happy to share our new paper on using AlphaEvolve for mathematical exploration at scale, written with Javier Gómez-Serrano, Terence Tao, and @GoogleDeepMind's Bogdan Georgiev. We tested it on 67 problems and documented all our successes and failures. 🧵 [image]

@azwagner_ Adam Zsolt Wagner

Context & Ripple Effects

This is a more systematic test of AlphaEvolve after DeepMind introduced it as an evolutionary coding agent for algorithm design and optimization the AlphaEvolve launch. It extends a research line that included FunSearch's result on the cap set problem and specialized mathematical-reasoning systems.

The reported hit rate matters because it records both successes and failures across a defined problem set, rather than relying on a single standout result. That offers a more concrete basis for judging where human-AI mathematical collaboration may be useful.

First-order effects

  • The participating researchers now have a documented evaluation of AlphaEvolve across 67 problems, including roughly 20 cases with improved solutions.
  • For DeepMind, the result supplies evidence that its evolutionary-agent approach can contribute beyond the geometry and Olympiad-oriented capabilities previously associated with AlphaGeometry and AlphaProof DeepMind's math-specialist model rollout.

Second-order effects

  • Mathematical AI teams face pressure to publish broader, failure-inclusive evaluations rather than foregrounding isolated discoveries, since comparative evidence is more useful to researchers choosing tools.
  • Researchers may increasingly use coding agents to generate and refine candidate approaches, while reserving human effort for selecting problems, checking results, and assessing significance.

Third-order effects

  • If replicated across other problem classes, mathematical AI could evolve from benchmark-focused solvers into a research workflow layer that expands the number of ideas experts can evaluate; the quality of verification and reporting will determine its practical value.
  • The field may shift toward auditable human-AI research processes, where reproducible evaluation sets and documented failures matter as much as individual novel solutions.

The trend: AI mathematics is moving from narrowly scored problem-solving systems toward evaluated agents designed to augment exploratory research workflows.

Discussion

  • @pushmeet Pushmeet Kohli on x
    (1) Our team at @GoogleDeepMind has been collaborating with Terence Tao and Javier Gómez-Serrano to use our AI agents (AlphaEvolve, AlphaProof, & Gemini Deep Think) for advancing Maths research. They find that AlphaEvolve can help discover new results across a range of problems.
  • @abigail_e_see Abigail See on x
    Really cool results on AI-assisted mathematics research, using AlphaEvolve to scalably discover constructions in a variety of mathematical disciplines!
  • @azwagner_ Adam Zsolt Wagner on x
    We tested AlphaEvolve on a wide-ranging portfolio of 67 mathematical problems in combinatorics, geometry and more. It found new constructions that improved upon the best-known results on about 20 of them, such as finding denser ways to pack 11 cubes into a larger cube. [image]
  • @g_leech_ Gavin Leech on x
    A gang of Google models attempt 67 solved and unsolved problems, aided by Terry Tao. As in past work they “just” involve improving bounds, mostly finite special cases (though not always!). Also solved IMO2025 P6 (but couldn't prove that it was optimal). https://terrytao.wordpress…
  • @abbas_mehrabian Abbas Mehrabian on x
    Incredible work! My DeepMind colleagues and prominent mathematicians collaborated to use AI to advance progress on over 20 mathematical problems. I was honoured to serve as a reviewer for this paper. Paper: https://arxiv.org/... Tao's blog post: https://terrytao.wordpress.com/ ..…
  • @richardcsuwandi Richard C. Suwandi on x
    LLM-guided evolutionary search is proving to be a powerful tool for mathematical discovery. On 67 math problems (mathematical analysis, combinatorics, geometry, number theory), AlphaEvolve was able to rediscover best-known solutions, find improved ones, and sometimes even [image]
  • @dmitryrybin1 Dmitry Rybin on x
    Terence Tao, Javier Gómez-Serrano, and DeepMind reported more extended experiments with AlphaEvolve + DeepThink + AlphaProof. Many new discoveries, including some general constructions that inspired T. Tao to write two new papers. We live in the wildest timeline [image]
  • @greghburnham Greg Burnham on x
    I had missed that AlphaEvolve is able to find the construction needed for the 2025 IMO P6. At least, when given this delightful hint... [image]