/
Navigation
Chronicles
Browse all articles
Explore
Semantic exploration
Research
Entity momentum
Nexus
Correlations & relationships
Story Arc
Topic evolution
Drift Map
Semantic trajectory animation
Posts
Analysis & commentary
Pulse API
Tech news intelligence API
Browse
Entities
Companies, people, products, technologies
Domains
Browse by publication source
Handles
Browse by social media handle
Detection
Concept Search
Semantic similarity search
High Impact Stories
Top coverage by position
Sentiment Analysis
Positive/negative coverage
Anomaly Detection
Unusual coverage patterns
Analysis
Rivalry Report
Compare two entities head-to-head
Semantic Pivots
Narrative discontinuities
Crisis Response
Event recovery patterns
Connected
Search: /
Command: ⌘K
Embeddings: large
TEXXR

Chronicles

The story behind the story

days · browse · Enter similar · o open

A look at the more challenging AI evaluations emerging in response to the rapid progress of models, including FrontierMath, Humanity's Last Exam, and RE-Bench

more interesting than it sounds! LinkedIn: Ross Dawson : The frontier of “evals”.  Evaluations comparing AI ahd human capabilities are evolving rapidly as AI rapidly leaves existing benchmarks in the dust. …

Time Tharin Pillay

Context & Ripple Effects

Public benchmarks were already losing discriminatory power: major AI companies had begun building internal tests as models approached or exceeded 90% accuracy on existing public evaluations internal benchmarks replaced saturated public tests. The new evaluation suite is therefore less a single scorecard than an attempt to restore meaningful separation among frontier models.

The subsequent release of Humanity’s Last Exam underscores how quickly benchmark design itself has become a moving target, alongside FrontierMath and RE-Bench.

First-order effects

  • Frontier-model developers and evaluators gain tougher instruments for distinguishing performance where older tests no longer do so reliably.
  • Claims of model progress face a higher evidentiary bar: comparisons increasingly depend on task sets designed to remain difficult for both models and humans.

Second-order effects

  • Labs that relied on familiar public-leaderboard results have stronger incentives to keep proprietary evaluations, making cross-company performance claims harder to audit.
  • Benchmark designers become a more consequential part of the AI ecosystem, because test quality shapes which model capabilities are visible and comparable.

Third-order effects

  • If rapid benchmark saturation persists, AI evaluation is likely to shift from static public tests toward a continual cycle of new, harder, and more specialized assessments.
  • That shift can make headline scores less durable as a common measure of progress, increasing the importance of evaluation methodology alongside the score itself.

The trend: Frontier AI measurement is moving from stable public benchmarks toward continuously refreshed evaluations built to track rapidly advancing models.

Discussion

  • @iseff.com Ian Sefferman on bluesky
    Interesting to think about the difficulty in creating *good* evals.  [embedded post]
  • @tharin_p @tharin_p on x
    My latest piece for @TIME contextualises o3's benchmark results with a look at the new wave of evals shaping the field, including @EpochAIResearch's FrontierMath, @METR_Evals's RE-Bench, and @scale_AI / @ai_risks's Humanity's Last Exam:
  • @tharin_p @tharin_p on x
    Merry Christmas everyone here is 2500 words on AI evals — more interesting than it sounds!