Q&A with mathematicians behind the “First Proof” experiment, which tests AI's mathematical competency on questions drawn from the authors' unpublished research
Large language models struggle to solve research-level math questions. It takes a human to measure just how poorly they perform.
Context & Ripple Effects
First Proof adds a harder evaluation layer to AI-for-mathematics coverage: rather than relying on established problems, it uses questions from researchers’ unpublished work and has mathematicians judge the results. That makes its reported weak performance a constraint on broad claims of mathematical competence.
The experiment arrives after specialized systems such as DeepMind’s math-reasoning models and FunSearch’s result on a longstanding pure-math problem demonstrated that AI can make meaningful contributions in bounded settings. It also provides a practical counterweight as newer reasoning models are described as increasingly useful for mathematics.
First-order effects
- AI labs and users get evidence that performance on research-level mathematics remains materially below what demonstrations on narrower or familiar tasks may imply.
- Mathematicians become essential evaluators: assessing whether an answer advances unpublished work cannot be delegated solely to model-generated verification.
Second-order effects
- Developers making math-capability claims face pressure to test on fresher, expert-curated questions, not only public benchmarks whose answers and solution styles may be familiar to models.
- Research teams using AI for mathematics are likely to treat it as an assistive tool requiring expert review, rather than a substitute for researchers on frontier problems.
Third-order effects
- If unpublished-work evaluations become widely adopted, frontier math benchmarks could shift from static scoreboards toward expert-run, continuously refreshed testing—raising the importance and cost of credible evaluation.
- The wider pattern is a separation between AI systems that can accelerate pieces of research and systems that can independently sustain reliable original reasoning; whether that gap narrows remains an empirical question.
The trend: AI mathematics is moving from headline demonstrations toward adversarial, expert-led evaluation of whether reasoning systems generalize to genuinely new research problems.