/
Navigation
Chronicles
Browse all articles
Explore
Semantic exploration
Research
Entity momentum
Nexus
Correlations & relationships
Story Arc
Topic evolution
Drift Map
Semantic trajectory animation
Posts
Analysis & commentary
Pulse API
Tech news intelligence API
Browse
Entities
Companies, people, products, technologies
Domains
Browse by publication source
Handles
Browse by social media handle
Detection
Concept Search
Semantic similarity search
High Impact Stories
Top coverage by position
Sentiment Analysis
Positive/negative coverage
Anomaly Detection
Unusual coverage patterns
Analysis
Rivalry Report
Compare two entities head-to-head
Semantic Pivots
Narrative discontinuities
Crisis Response
Event recovery patterns
Connected
Search: /
Command: ⌘K
Embeddings: large
TEXXR

Chronicles

The story behind the story

days · browse · Enter similar · o open

Q&A with mathematicians behind the “First Proof” experiment, which tests AI's mathematical competence on questions drawn from the authors' unpublished research

Large language models struggle to solve research-level math questions.  It takes a human to measure just how poorly they perform.

New York Times Siobhan Roberts

Context & Ripple Effects

AI mathematics claims have moved from narrow, specialized systems—such as DeepMind’s math-reasoning models—to broader arguments that newer reasoning models are more useful to mathematicians. First Proof supplies a deliberately tougher check: questions rooted in work not yet public.

That matters because prior reports of AI research progress, including FunSearch’s result on a long-standing mathematics problem, do not establish reliable performance across unfamiliar research questions. The experiment centers mathematicians’ own evaluation of that gap.

First-order effects

  • First Proof gives mathematicians a research-level, unpublished-work benchmark on which current large language models perform poorly, limiting claims that they can independently handle this class of mathematical work.
  • AI developers and users seeking mathematical assistance must distinguish demonstrated help on selected tasks from dependable solutions to novel research problems.

Second-order effects

  • Model developers face pressure to test reasoning systems against less-contaminated, expert-authored problems rather than rely on benchmarks whose answers may be widely available.
  • Mathematics researchers and AI labs are likely to place more value on human verification and domain-specific evaluation when deciding where models can accelerate research workflows.

Third-order effects

  • If unpublished-problem testing becomes more common, AI progress in mathematics will be judged less by isolated breakthroughs and more by robustness on genuinely novel work.
  • The field may develop a clearer division between AI as a research aid and AI as an autonomous mathematical reasoner; the boundary will depend on reproducible expert evaluation, not model output alone.

The trend: AI mathematics is shifting from headline-grabbing demonstrations toward expert-designed tests of whether reasoning systems generalize beyond known answers.

Discussion

  • @jugander Johan Ugander on bluesky
    Ten math problems with proofs known the authors.  Proofs are encrypted until Feb 13.  For all problems, authors claim both AI-based literature searches and zero-shot attempts at proofs failed.  If you want to take a crack, you have until next Friday (2/13)!
  • @javifields Javier Campos on bluesky
    To assess the ability of current AI systems to correctly answer research-level mathematics questions, we share a set of ten math questions which have arisen naturally in the research process of the authors.  —  arxiv.org/abs/2602.05192
  • r/math r on reddit
    These Mathematicians Are Putting A.I. to the Test