/
Navigation
Chronicles
Browse all articles
Explore
Semantic exploration
Research
Entity momentum
Nexus
Correlations & relationships
Story Arc
Topic evolution
Drift Map
Semantic trajectory animation
Posts
Analysis & commentary
Pulse API
Tech news intelligence API
Browse
Entities
Companies, people, products, technologies
Domains
Browse by publication source
Handles
Browse by social media handle
Detection
Concept Search
Semantic similarity search
High Impact Stories
Top coverage by position
Sentiment Analysis
Positive/negative coverage
Anomaly Detection
Unusual coverage patterns
Analysis
Rivalry Report
Compare two entities head-to-head
Semantic Pivots
Narrative discontinuities
Crisis Response
Event recovery patterns
Connected
Search: /
Command: ⌘K
Embeddings: large
TEXXR

Chronicles

The story behind the story

days · browse · Enter similar · o open

Q&A with mathematicians behind the “First Proof” experiment, which tests AI's mathematical competency on questions drawn from the authors' unpublished research

Large language models struggle to solve research-level math questions.  It takes a human to measure just how poorly they perform.

New York Times Siobhan Roberts

Context & Ripple Effects

First Proof adds a harder evaluation layer to AI-for-mathematics coverage: rather than relying on established problems, it uses questions from researchers’ unpublished work and has mathematicians judge the results. That makes its reported weak performance a constraint on broad claims of mathematical competence.

The experiment arrives after specialized systems such as DeepMind’s math-reasoning models and FunSearch’s result on a longstanding pure-math problem demonstrated that AI can make meaningful contributions in bounded settings. It also provides a practical counterweight as newer reasoning models are described as increasingly useful for mathematics.

First-order effects

  • AI labs and users get evidence that performance on research-level mathematics remains materially below what demonstrations on narrower or familiar tasks may imply.
  • Mathematicians become essential evaluators: assessing whether an answer advances unpublished work cannot be delegated solely to model-generated verification.

Second-order effects

  • Developers making math-capability claims face pressure to test on fresher, expert-curated questions, not only public benchmarks whose answers and solution styles may be familiar to models.
  • Research teams using AI for mathematics are likely to treat it as an assistive tool requiring expert review, rather than a substitute for researchers on frontier problems.

Third-order effects

  • If unpublished-work evaluations become widely adopted, frontier math benchmarks could shift from static scoreboards toward expert-run, continuously refreshed testing—raising the importance and cost of credible evaluation.
  • The wider pattern is a separation between AI systems that can accelerate pieces of research and systems that can independently sustain reliable original reasoning; whether that gap narrows remains an empirical question.

The trend: AI mathematics is moving from headline demonstrations toward adversarial, expert-led evaluation of whether reasoning systems generalize to genuinely new research problems.

Discussion

  • @jugander Johan Ugander on bluesky
    Ten math problems with proofs known the authors.  Proofs are encrypted until Feb 13.  For all problems, authors claim both AI-based literature searches and zero-shot attempts at proofs failed.  If you want to take a crack, you have until next Friday (2/13)!
  • @javifields Javier Campos on bluesky
    To assess the ability of current AI systems to correctly answer research-level mathematics questions, we share a set of ten math questions which have arisen naturally in the research process of the authors.  —  arxiv.org/abs/2602.05192
  • r/math r on reddit
    These Mathematicians Are Putting A.I. to the Test