At the International Mathematical Olympiad, 26 students got higher scores than DeepMind and OpenAI models, possibly the last time humans beat AI at the exam
Google DeepMind and OpenAI won gold medals at the math Olympics—but these American teenagers still got higher scores.
Context & Ripple Effects
DeepMind’s path to this result began with AlphaGeometry approaching gold-medalist geometry performance and later expanded to AlphaGeometry2’s reported performance across past Olympiad geometry questions. OpenAI’s experimental reasoning model had also been reported at gold-medal level shortly before this year’s contest in a researcher’s account of its Olympiad result.
The new distinction is not whether the models cleared the gold threshold, but that 26 students still exceeded them on the same exam. That makes the Olympiad a sharper measure of the remaining gap at the top end of mathematical reasoning.
First-order effects
- DeepMind and OpenAI gain a high-profile validation of their models’ contest-math capabilities, while the highest-scoring students retain a measurable lead on this examination.
- The International Mathematical Olympiad becomes a more nuanced benchmark: gold-medal status alone no longer captures separation among elite performers and leading models.
Second-order effects
- AI labs will face pressure to report more than medal-equivalent outcomes—such as score distributions and performance against the very top competitors—when presenting reasoning progress.
- Researchers and evaluators may put greater weight on difficult, end-to-end proof tasks rather than isolated geometry results, despite AlphaGeometry2’s strong historical geometry showing.
Third-order effects
- If leading models continue to close this gap, elite competition mathematics may shift from a headline benchmark toward a diagnostic for specific weaknesses in reasoning, verification, and problem selection.
- The broader value of mathematical benchmarks will increasingly depend on credible evaluation design: models reaching a threshold is different from reliably outperforming the strongest humans.
The trend: Frontier AI competition is moving from threshold-based benchmark wins toward scrutiny of consistency and performance at the extreme human frontier.