/
Navigation
Chronicles
Browse all articles
Explore
Semantic exploration
Research
Entity momentum
Nexus
Correlations & relationships
Story Arc
Topic evolution
Drift Map
Semantic trajectory animation
Posts
Analysis & commentary
Pulse API
Tech news intelligence API
Browse
Entities
Companies, people, products, technologies
Domains
Browse by publication source
Handles
Browse by social media handle
Detection
Concept Search
Semantic similarity search
High Impact Stories
Top coverage by position
Sentiment Analysis
Positive/negative coverage
Anomaly Detection
Unusual coverage patterns
Analysis
Rivalry Report
Compare two entities head-to-head
Semantic Pivots
Narrative discontinuities
Crisis Response
Event recovery patterns
Connected
Search: /
Command: ⌘K
Embeddings: large
TEXXR

Chronicles

The story behind the story

days · browse · Enter similar · o open

A look at the more challenging AI evaluations emerging in response to the rapid progress of models, including FrontierMath, Humanity's Last Exam, and RE-Bench

Despite their expertise, AI developers don't always know what their most advanced systems are capable of—at least, not at first. X: @tharin_p and @tharin_p X: @tharin_p : My latest piece for @TIME contextualises o3's benchmark results with a look at the new wave of evals shaping the field, including @EpochAIResearch's FrontierMath, @METR_Evals's RE-Bench, and @scale_AI / @ai_risks's Humanity's Last Exam: @tharin_p : Merry Christmas everyone here is 2500 words on AI evals — more interesting than it sounds!

Time Tharin Pillay

Context & Ripple Effects

Recent results such as o3's 87.5% score on ARC Prize's semi-private evaluation have underscored how quickly established tests can lose discriminatory power. Coverage also showed major labs building their own internal benchmarks as public measures approach ceiling performance.

FrontierMath, RE-Bench, and Humanity's Last Exam form part of the response: evaluations designed to reveal capabilities and high-end behavior that developers may not identify immediately from routine testing.

First-order effects

  • Model developers gain harder external tests for probing advanced systems, while evaluation groups become more central sources of evidence about capability limits and behavior.
  • Benchmark results become less easily summarized by legacy test scores, particularly for systems such as o3 that are presented as reasoning-oriented models.

Second-order effects

  • Labs that rely on internal benchmarks face pressure to demonstrate performance on independently developed evaluations, rather than treating proprietary tests as sufficient evidence.
  • More demanding evaluations raise the cost and complexity of model assessment, shifting attention from headline accuracy rates toward the conditions, task types, and compute configurations behind results.

Third-order effects

  • If capability gains continue to outpace existing tests, evaluation will become an ongoing measurement discipline rather than a one-time pre-release comparison—an institutional role shared by labs and specialist evaluators.
  • The gap between what developers initially observe and what models can do strengthens the case for operational assurance, though the corpus does not establish a common standard for it yet.

The trend: Frontier AI is moving from broad benchmark competition toward continuous, adversarial evaluation of increasingly capable reasoning systems.

Discussion

  • @tharin_p @tharin_p on x
    My latest piece for @TIME contextualises o3's benchmark results with a look at the new wave of evals shaping the field, including @EpochAIResearch's FrontierMath, @METR_Evals's RE-Bench, and @scale_AI / @ai_risks's Humanity's Last Exam:
  • @tharin_p @tharin_p on x
    Merry Christmas everyone here is 2500 words on AI evals — more interesting than it sounds!