Q&A with mathematicians behind the “First Proof” experiment, which tests AI's mathematical competence on questions drawn from the authors' unpublished research
Large language models struggle to solve research-level math questions. It takes a human to measure just how poorly they perform.
Context & Ripple Effects
AI mathematics claims have moved from narrow, specialized systems—such as DeepMind’s math-reasoning models—to broader arguments that newer reasoning models are more useful to mathematicians. First Proof supplies a deliberately tougher check: questions rooted in work not yet public.
That matters because prior reports of AI research progress, including FunSearch’s result on a long-standing mathematics problem, do not establish reliable performance across unfamiliar research questions. The experiment centers mathematicians’ own evaluation of that gap.
First-order effects
- First Proof gives mathematicians a research-level, unpublished-work benchmark on which current large language models perform poorly, limiting claims that they can independently handle this class of mathematical work.
- AI developers and users seeking mathematical assistance must distinguish demonstrated help on selected tasks from dependable solutions to novel research problems.
Second-order effects
- Model developers face pressure to test reasoning systems against less-contaminated, expert-authored problems rather than rely on benchmarks whose answers may be widely available.
- Mathematics researchers and AI labs are likely to place more value on human verification and domain-specific evaluation when deciding where models can accelerate research workflows.
Third-order effects
- If unpublished-problem testing becomes more common, AI progress in mathematics will be judged less by isolated breakthroughs and more by robustness on genuinely novel work.
- The field may develop a clearer division between AI as a research aid and AI as an autonomous mathematical reasoner; the boundary will depend on reproducible expert evaluation, not model output alone.
The trend: AI mathematics is shifting from headline-grabbing demonstrations toward expert-designed tests of whether reasoning systems generalize beyond known answers.