/
Navigation
Chronicles
Browse all articles
Explore
Semantic exploration
Research
Entity momentum
Nexus
Correlations & relationships
Story Arc
Topic evolution
Drift Map
Semantic trajectory animation
Posts
Analysis & commentary
Pulse API
Tech news intelligence API
Browse
Entities
Companies, people, products, technologies
Domains
Browse by publication source
Handles
Browse by social media handle
Detection
Concept Search
Semantic similarity search
High Impact Stories
Top coverage by position
Sentiment Analysis
Positive/negative coverage
Anomaly Detection
Unusual coverage patterns
Analysis
Rivalry Report
Compare two entities head-to-head
Semantic Pivots
Narrative discontinuities
Crisis Response
Event recovery patterns
Connected
Search: /
Command: ⌘K
Embeddings: large
TEXXR

Chronicles

The story behind the story

days · browse · Enter similar · o open

CAIS and Scale AI release “Humanity's Last Exam”, which they claim is the hardest- AI test yet, consisting of ~3,000 multiple-choice and short answer questions

If you're looking for a new reason to be nervous about artificial intelligence, try this: Some of the smartest humans …

New York Times Kevin Roose

Context & Ripple Effects

Harder evaluations were already emerging as rapid model progress made established tests less discriminating, including the set of new benchmarks surveyed in coverage of tougher AI evaluations. CAIS and Scale AI now add a large, public challenge intended to raise that bar.

The release matters because benchmark design increasingly shapes how model capability claims are compared. A test spanning multiple-choice and short-answer formats can become a common reference point, but only if researchers and developers treat its results as meaningful rather than a single headline score.

First-order effects

  • CAIS and Scale AI gain a new vehicle for measuring and publicizing model performance, while model developers receive a harder external test on which to assess their systems.
  • Researchers and buyers evaluating frontier models get another standardized signal, though the organizers' claim of exceptional difficulty still depends on how models perform and how the benchmark is used.

Second-order effects

  • Competing AI labs and benchmark creators face pressure to report against more demanding evaluations or explain why their preferred tests better capture useful capabilities.
  • As high-profile tests become targets for optimization, developers will have stronger incentives to distinguish genuine generalization from performance tuned to a known benchmark.

Third-order effects

  • If this pattern persists, AI evaluation will become an ongoing arms race: new tests will be needed as older ones lose their ability to separate leading models.
  • The durable shift is toward operational assurance based on portfolios of evaluations rather than a single score, with credibility depending on test design, freshness, and resistance to targeted training.

The trend: Frontier AI progress is driving a shift from static benchmark leaderboards toward continually refreshed, harder evaluations meant to test whether capabilities generalize.

Discussion

  • @kevinroose Kevin Roose on x
    It must be comforting, on some level, to believe that AI progress is hitting a wall. But the reality is that the industry is scrambling to design new tests hard enough to stump AI models, because most of the existing tests are getting beaten. Full column (free link):
  • @kevinroose Kevin Roose on x
    New by me: There's a new AI evaluation called “Humanity's Last Exam,” with ~3,000 questions drawn from leading academics and experts. It's the hardest AI test ever — today, no model gets above 10% — but researchers expect 50% scores by the end of the year. Anyway, carry on! [imag…