DeepMind says its AlphaGeometry2 model solved 84% of International Math Olympiad's geometry problems from the last 25 years, surpassing average gold medalists
The International Math Olympiad is becoming a recurring stress test for specialized AI reasoning systems, making comparative performance on its geometry archive a visible measure of progress rather than a one-off model demo.
First-order effects
DeepMind gains a stronger benchmark claim for AlphaGeometry2: across the stated historical geometry set, its reported performance exceeds the average gold medalist benchmark.
The result raises the bar for how DeepMind’s math-focused models are assessed, shifting attention from near-human performance to consistency across a long-running problem archive.
Second-order effects
Other frontier-model developers will face pressure to demonstrate comparable performance on rigorous, independently recognizable reasoning benchmarks rather than broad claims of math ability.
Teams building math-reasoning systems may increasingly separate geometry, proof, and general reasoning capabilities, reflecting DeepMind’s paired AlphaProof/AlphaGeometry approach.
Third-order effects
If repeatable across new problems and external evaluations, Olympiad-style testing could become a more consequential proving ground for AI systems that claim formal reasoning competence.
The broader shift is from general-purpose model comparisons toward domain-specific systems whose value rests on verifiable outputs and benchmark design.
The trend: AI labs are turning formal mathematics into a high-visibility benchmark for measuring whether specialized reasoning systems are becoming reliably capable rather than merely fluent.
@umesh_ai Not very well given that at least the public version doesn't support images. Even prior versions of public LLMs don't have the visual fidelity to even interpret these problems. Likely 0%.
Very interesting contrast here with the people checking LLM results on my derivatives interview questions. Result is always either LLM waves hands and says various boilerplate true statements that do not answer the question, or LLM confidently states precise wrong answer
Another remarkable AI milestone! Congratulations to @GoogleDeepMind for the incredible achievement of Alpha Geometry 2! I believe that when AI begins to formulate questions humans can't solve, we will have reached Artificial Super Intelligence!
If you read the paper, it's an impressive algorithm but hardly AGI. They're using an LLM to turn the problem statement into a pseudocode-like representation, then generating a diagram, and the DDAR algorithm iterates over every point trying to deduce facts from postulates.
HUGE: Google's AI just solved 84% of the International Math Olympiad (IMO) problems from 2000-24 with Alpha Geometry 2! These are math problems most professors couldn't solve. [image]