/
Navigation
Chronicles
Browse all articles
Explore
Semantic exploration
Research
Entity momentum
Nexus
Correlations & relationships
Story Arc
Topic evolution
Drift Map
Semantic trajectory animation
Posts
Analysis & commentary
Pulse API
Tech news intelligence API
Browse
Entities
Companies, people, products, technologies
Domains
Browse by publication source
Handles
Browse by social media handle
Detection
Concept Search
Semantic similarity search
High Impact Stories
Top coverage by position
Sentiment Analysis
Positive/negative coverage
Anomaly Detection
Unusual coverage patterns
Analysis
Rivalry Report
Compare two entities head-to-head
Semantic Pivots
Narrative discontinuities
Crisis Response
Event recovery patterns
Connected
Search: /
Command: ⌘K
Embeddings: large
TEXXR

Chronicles

The story behind the story

days · browse · Enter similar · o open

A look at the AI nonprofit METR, whose time-horizon metrics are used by AI researchers and Wall Street investors to track the rapid development of AI systems

New York Times Kevin Roose

Context & Ripple Effects

METR’s task-time-horizon measure is moving beyond model evaluation into an investor-facing indicator of AI progress. That expands the audience for evaluations at a time when coverage has focused on tougher tests for rapidly improving models and on the competitive stakes for major platforms and foundation-model makers.

The reported roughly four-to-four-and-a-half-month doubling of the horizon at which systems reach 50% task reliability gives researchers and markets a shared, if necessarily partial, way to describe the pace of capability gains.

First-order effects

  • AI researchers and Wall Street investors gain a common metric for tracking changes in model reliability over longer tasks, making METR more consequential in how progress is communicated.
  • The reported rate of improvement raises the practical importance of testing whether models can complete multi-step work reliably, rather than relying only on narrower benchmark results.

Second-order effects

  • Model developers face stronger incentives to demonstrate gains on evaluations that connect capability to economically meaningful task duration, not just headline benchmark scores.
  • Investors may increasingly use evaluation results as inputs to judgments about the timing and scope of AI-driven productivity or revenue changes, amplifying scrutiny of how those metrics are constructed and interpreted.

Third-order effects

  • If time-horizon measures become widely accepted, independent evaluation groups could become part of the market infrastructure that translates frontier-model progress into deployment and investment decisions.
  • The broader shift is toward institutionalized AI measurement: benchmarks may shape commercial expectations, but their influence will depend on whether they remain robust across real-world tasks and model changes.

The trend: AI capability evaluation is becoming a shared layer of research, product strategy, and financial market analysis as models are assessed on longer and more reliable task execution.

Discussion

  • @htbroadley Thomas Broadley on x
    My team at METR is making the graph in the photo! It is so wild for that to be in NYT. We are hiring :-)
  • @chrispainteryup Chris Painter on x
    Cool profile of METR's work in the NYT today! I particularly like this from @ajeya_cotra: “METR is an organization that asks... what we think would be most valuable for the world to know about A.I. and its risks, and then the answers are what they are.” https://www.nytimes.com/..…
  • @kevinroose Kevin Roose on x
    New column: I went to visit @METR_Evals, the 30-person AI nonprofit that makes the Most Important Chart in the World. I learned a lot, but the most striking thing was how soon some of them think AI R&D could be fully automated. (This year!) https://www.nytimes.com/...
  • Rasmus Faber-Espensen Rasmus Faber-Espensen on linkedin
    I've worked with a lot of great people over the years, but I don't think I have ever been somewhere where I daily encounter so many scarily intelligent people …