/
Navigation
Chronicles
Browse all articles
Explore
Semantic exploration
Research
Entity momentum
Nexus
Correlations & relationships
Story Arc
Topic evolution
Drift Map
Semantic trajectory animation
Posts
Analysis & commentary
Pulse API
Tech news intelligence API
Browse
Entities
Companies, people, products, technologies
Domains
Browse by publication source
Handles
Browse by social media handle
Detection
Concept Search
Semantic similarity search
High Impact Stories
Top coverage by position
Sentiment Analysis
Positive/negative coverage
Anomaly Detection
Unusual coverage patterns
Analysis
Rivalry Report
Compare two entities head-to-head
Semantic Pivots
Narrative discontinuities
Crisis Response
Event recovery patterns
Connected
Search: /
Command: ⌘K
Embeddings: large
TEXXR

Chronicles

The story behind the story

days · browse · Enter similar · o open

A look at the AI nonprofit METR, whose time-horizon metrics are used by AI researchers and Wall Street investors to track the rapid development of AI systems

A chart created by METR, a nonprofit A.I. organization, has become an industrywide obsession as it measures the rapid development of big A.I. systems.

New York Times Kevin Roose

Context & Ripple Effects

METR’s time-horizon chart has moved beyond a research artifact: researchers and Wall Street investors use it as a common yardstick for how quickly AI systems are extending the duration of tasks they can complete reliably. Its reported 50%-reliability horizon is roughly doubling every four to four-and-a-half months.

The coverage sits alongside a broader shift toward tougher AI evaluations, including FrontierMath, Humanity’s Last Exam, and RE-Bench. As models advance, evaluation is becoming more consequential both for judging capability claims and for translating technical progress into business expectations.

First-order effects

  • METR gains outsized influence over how researchers and investors interpret frontier-model progress, because its metric offers a compact benchmark for comparing changes in useful task completion.
  • Model developers and AI adopters face greater pressure to demonstrate reliability over longer, real-world task sequences rather than relying on narrow or easily saturated benchmarks.

Second-order effects

  • Investors may use time-horizon results to reassess which AI applications are nearing practical automation, shifting attention from model novelty toward the length and reliability of work systems can delegate.
  • Benchmark designers and labs are pushed to build harder, more representative evaluations; otherwise widely watched metrics risk becoming less informative as systems adapt to them.

Third-order effects

  • If time-horizon measurement remains credible, AI capability assessment could become a more institutionalized input to capital allocation and enterprise deployment decisions, not solely a research exercise.
  • The key constraint will increasingly be measurement quality: progress narratives will depend on whether independent evaluations capture dependable performance on meaningful work, rather than isolated test gains.

The trend: AI evaluation is evolving from a technical scoreboard into market infrastructure for estimating when increasingly capable systems can perform economically useful work.

Discussion

  • @chrispainteryup Chris Painter on x
    Cool profile of METR's work in the NYT today! I particularly like this from @ajeya_cotra: “METR is an organization that asks... what we think would be most valuable for the world to know about A.I. and its risks, and then the answers are what they are.” https://www.nytimes.com/..…
  • @htbroadley Thomas Broadley on x
    My team at METR is making the graph in the photo! It is so wild for that to be in NYT. We are hiring :-)
  • @kevinroose Kevin Roose on x
    New column: I went to visit @METR_Evals, the 30-person AI nonprofit that makes the Most Important Chart in the World. I learned a lot, but the most striking thing was how soon some of them think AI R&D could be fully automated. (This year!) https://www.nytimes.com/...
  • Rasmus Faber-Espensen Rasmus Faber-Espensen on linkedin
    I've worked with a lot of great people over the years, but I don't think I have ever been somewhere where I daily encounter so many scarily intelligent people …
  • r/technology r on reddit
    NYT- How Do You Measure an A.I. Boom?