/
Navigation
Chronicles
Browse all articles
Explore
Semantic exploration
Research
Entity momentum
Nexus
Correlations & relationships
Story Arc
Topic evolution
Drift Map
Semantic trajectory animation
Posts
Analysis & commentary
Pulse API
Tech news intelligence API
Browse
Entities
Companies, people, products, technologies
Domains
Browse by publication source
Handles
Browse by social media handle
Detection
Concept Search
Semantic similarity search
High Impact Stories
Top coverage by position
Sentiment Analysis
Positive/negative coverage
Anomaly Detection
Unusual coverage patterns
Analysis
Rivalry Report
Compare two entities head-to-head
Semantic Pivots
Narrative discontinuities
Crisis Response
Event recovery patterns
Connected
Search: /
Command: ⌘K
Embeddings: large
TEXXR

Chronicles

The story behind the story

← → days · ↑ ↓ browse · Enter similar · o open

METR: Claude Opus 4.5 has a 50% task completion time horizon of about 4 hours and 49 minutes, more than double that of Claude Opus 4 released earlier this year

just careful, meticulous rigor. Nikola Jurkovic / @nikolaj2030 : This result updates me towards 4 month doubling times being my median estimate for the next two years. That means by EOY 2026 the time horizon will be 40 hours, and by EOY 2027 it will be 320 hours. @aidigest_ : Opus 4.5 puts the world roughly back on track for the red line 😬 Every ~4 months, the length of coding tasks AI agents can perform (compared to human professionals) *doubles* More context on this finding in @METR_Evals thread https://x.com/... [image] James Moore / @jamesdmoore614 : https://www.youtube.com/... Mathematical proof behind the “100x improvement” prediction shows that reliability is scaling log-linearly, which is exactly what is needed to turn AI from a “chatbot” into the “autonomous worker” the Moonshots team predicted. @PeterDiamandis Shakeel / @shakeelhashim : Yesterday I wrote that the performance of AI models on METRs time horizon tasks had increased by 4x this year. Turns out it's actually 7x. Makes my point from yesterday even more valid, I think — despite all the talk of AI progress slowing down this year, things have continued [image] Frank Downing / @downingark : Expanding time horizons on this eval represent AI's increasing ability to reliably deliver leverage on human labor. As basis point improvements on PHD level quizzes seem less and less representative of real world performance, the METR study is increasingly relevant. David Shapiro / @daveshapi : I downloaded the latest data and repeated my previous exercises from August of this year to determine the current “best fit” model. TLDR: we are looking at a log quadratic fit. When we forecast this out, we're looking at “human equivalent task length” quickly stretching into [image] Peter Wildeford / @peterwildeford : Recent METR data shows ~4.4mo doubling for 50% reliability, with frontier models now completing median tasks up to ~5 hours. Which is crazy! 80% reliability fits similar ~4.5mo, but w/ more variance. Reliability gains may be plateauing somewhat? Worth monitoring. [image] Dean W. Ball / @deanwball : I wonder if the average software engineer, drawn from the worldwide population, achieves a >50% success rate on tasks that take METR's sample engineers ~4.5 hours. Ryan Greenblatt / @ryanpgreenblatt : I've updated towards a higher chance of much faster than trend AI progress in 2026. I also generally updated towards (somewhat) faster doubling times in the next 2 years. I was expecting ~170 day doubling times (5.7 months) and now I expect ~150 day doubling times (5 months). Bluesky: Jeremy Diamond / @dmnd.me : There's a bur after you get past the toy/novelty uses where the learning curve gets steep fast.  —  I don't see a lot of discussion about how much more effort and scaffolding you need to put in to enable these products to run this long and actually produce good output metr.org/blog/2025-03... Ethan Mollick / @emollick : Since people are asking about the “exponential growth” bit - AI has a jagged frontier but on most hard careful measures of AI ability we are seeing rapid exponential gains until benchmarks get saturated.  —  Take METR long tasks: metr.org/blog/2025-03...  [image] Alex Hern / @hern : METR has updated its LLM time-horizon chart for Claude 4.5 Sonnet and we're still above the trendline.  Slowdown?  What slowdown? metr.org/blog/2025-03...  [image]

@metr_evals

Context & Ripple Effects

METR's result adds a model-specific data point to a contested picture of coding-agent progress. Earlier coverage found that experienced developers using AI tools could be slower despite feeling faster, while a later assessment identified a “messiness tax” on lengthy real-world software tasks.

The result also arrives as Anthropic reported that its own employees use Claude heavily for debugging and code understanding, with self-reported productivity gains. That makes a longer task horizon relevant not just as a benchmark result but as a measure of how much work can plausibly be delegated before human review.

First-order effects

  • Claude Opus 4.5 can be evaluated against substantially longer coding tasks at the 50% completion threshold than Claude Opus 4, strengthening Anthropic's case for agentic development workflows.
  • Engineering teams gain a more concrete boundary for piloting longer-running Claude assignments, but the 50% metric still implies that verification and recovery remain integral to deployment.

Second-order effects

  • Competing model providers and coding-tool vendors face pressure to demonstrate reliability on longer, end-to-end tasks rather than emphasize short benchmark wins.
  • Buyers will increasingly evaluate AI coding tools by useful completed work and oversight burden; the earlier developer productivity study shows why perceived speed alone is an insufficient procurement metric.

Third-order effects

  • If task horizons continue extending while reliability improves, software work may shift from tool-assisted coding toward workflows in which people scope, supervise, and validate multi-hour agent runs.
  • The key competitive constraint would become dependable execution in messy production environments—not merely model capability—raising the value of evaluation, observability, and human-review infrastructure.

The trend: AI coding is moving toward longer-duration agent execution, with real-world reliability and cost per completed task becoming the decisive adoption measures.

Discussion

  • @idavidrein David Rein on x
    While I have a lot of uncertainty about this result (look at the confidence intervals!!), I do think this is just a clear example of what being on an exponential trend means, in terms of what it *feels like*. You gotta price this sort of thing in
  • @ninoyako93 @ninoyako93 on x
    The expontential trends for AI capabilities are accelerating, pushing progress into super-exponential trajectories. The next few years will be so transforming, people wont be able recognize the world they live in anymore.
  • @tyler_m_john Tyler John on x
    The long awaited Opus results. Opus breaks the task suite, which doesn't have long enough tasks for it to do. METR can't confidently rule out a 20 hour time horizon, which would about a year ahead of schedule. Epoch predicts Gemini 3 will be even better.
  • @liron Liron Shapira on x
    This AGI capability scale didn't even exist before 2020, yet whatever tech cascade it's measuring is incredibly robust, and shows no sign of S curving before blowing past human level. This is the WaitButWhy train whizzing past human station scenario.
  • @spoonedher @spoonedher on x
    few things to keep in mind 1. the confidence interval is insane and should raise suspicion 2. this is the 50% time horizon. super exponential relies on the 80% which is paltry 3. their sample size is really small and therefore exploitable through rlvr
  • @idavidrein David Rein on x
    Happy to have this out, there was a bunch of work that went into validating this result that I was really impressed by—just careful, meticulous rigor.
  • @nikolaj2030 Nikola Jurkovic on x
    This result updates me towards 4 month doubling times being my median estimate for the next two years. That means by EOY 2026 the time horizon will be 40 hours, and by EOY 2027 it will be 320 hours.
  • @aidigest_ @aidigest_ on x
    Opus 4.5 puts the world roughly back on track for the red line 😬 Every ~4 months, the length of coding tasks AI agents can perform (compared to human professionals) *doubles* More context on this finding in @METR_Evals thread https://x.com/... [image]
  • @jamesdmoore614 James Moore on x
    https://www.youtube.com/... Mathematical proof behind the “100x improvement” prediction shows that reliability is scaling log-linearly, which is exactly what is needed to turn AI from a “chatbot” into the “autonomous worker” the Moonshots team predicted. @PeterDiamandis
  • @shakeelhashim Shakeel on x
    Yesterday I wrote that the performance of AI models on METRs time horizon tasks had increased by 4x this year. Turns out it's actually 7x. Makes my point from yesterday even more valid, I think — despite all the talk of AI progress slowing down this year, things have continued [i…
  • @downingark Frank Downing on x
    Expanding time horizons on this eval represent AI's increasing ability to reliably deliver leverage on human labor. As basis point improvements on PHD level quizzes seem less and less representative of real world performance, the METR study is increasingly relevant.
  • @daveshapi David Shapiro on x
    I downloaded the latest data and repeated my previous exercises from August of this year to determine the current “best fit” model. TLDR: we are looking at a log quadratic fit. When we forecast this out, we're looking at “human equivalent task length” quickly stretching into [ima…
  • @peterwildeford Peter Wildeford on x
    Recent METR data shows ~4.4mo doubling for 50% reliability, with frontier models now completing median tasks up to ~5 hours. Which is crazy! 80% reliability fits similar ~4.5mo, but w/ more variance. Reliability gains may be plateauing somewhat? Worth monitoring. [image]
  • @deanwball Dean W. Ball on x
    I wonder if the average software engineer, drawn from the worldwide population, achieves a >50% success rate on tasks that take METR's sample engineers ~4.5 hours.
  • @ryanpgreenblatt Ryan Greenblatt on x
    I've updated towards a higher chance of much faster than trend AI progress in 2026. I also generally updated towards (somewhat) faster doubling times in the next 2 years. I was expecting ~170 day doubling times (5.7 months) and now I expect ~150 day doubling times (5 months).
  • @dmnd.me Jeremy Diamond on bluesky
    There's a bur after you get past the toy/novelty uses where the learning curve gets steep fast.  —  I don't see a lot of discussion about how much more effort and scaffolding you need to put in to enable these products to run this long and actually produce good output metr.org/bl…
  • @emollick Ethan Mollick on bluesky
    Since people are asking about the “exponential growth” bit - AI has a jagged frontier but on most hard careful measures of AI ability we are seeing rapid exponential gains until benchmarks get saturated.  —  Take METR long tasks: metr.org/blog/2025-03...  [image]
  • @hern Alex Hern on bluesky
    METR has updated its LLM time-horizon chart for Claude 4.5 Sonnet and we're still above the trendline.  Slowdown?  What slowdown? metr.org/blog/2025-03...  [image]