METR: Claude Opus 4.5 has a 50% task completion time horizon of about 4 hours and 49 minutes, more than double that of Claude Opus 4 released earlier this year
just careful, meticulous rigor. Nikola Jurkovic / @nikolaj2030 : This result updates me towards 4 month doubling times being my median estimate for the next two years. That means by EOY 2026 the time horizon will be 40 hours, and by EOY 2027 it will be 320 hours. @aidigest_ : Opus 4.5 puts the world roughly back on track for the red line 😬 Every ~4 months, the length of coding tasks AI agents can perform (compared to human professionals) *doubles* More context on this finding in @METR_Evals thread https://x.com/... [image] James Moore / @jamesdmoore614 : https://www.youtube.com/... Mathematical proof behind the “100x improvement” prediction shows that reliability is scaling log-linearly, which is exactly what is needed to turn AI from a “chatbot” into the “autonomous worker” the Moonshots team predicted. @PeterDiamandis Shakeel / @shakeelhashim : Yesterday I wrote that the performance of AI models on METRs time horizon tasks had increased by 4x this year. Turns out it's actually 7x. Makes my point from yesterday even more valid, I think — despite all the talk of AI progress slowing down this year, things have continued [image] Frank Downing / @downingark : Expanding time horizons on this eval represent AI's increasing ability to reliably deliver leverage on human labor. As basis point improvements on PHD level quizzes seem less and less representative of real world performance, the METR study is increasingly relevant. David Shapiro / @daveshapi : I downloaded the latest data and repeated my previous exercises from August of this year to determine the current “best fit” model. TLDR: we are looking at a log quadratic fit. When we forecast this out, we're looking at “human equivalent task length” quickly stretching into [image] Peter Wildeford / @peterwildeford : Recent METR data shows ~4.4mo doubling for 50% reliability, with frontier models now completing median tasks up to ~5 hours. Which is crazy! 80% reliability fits similar ~4.5mo, but w/ more variance. Reliability gains may be plateauing somewhat? Worth monitoring. [image] Dean W. Ball / @deanwball : I wonder if the average software engineer, drawn from the worldwide population, achieves a >50% success rate on tasks that take METR's sample engineers ~4.5 hours. Ryan Greenblatt / @ryanpgreenblatt : I've updated towards a higher chance of much faster than trend AI progress in 2026. I also generally updated towards (somewhat) faster doubling times in the next 2 years. I was expecting ~170 day doubling times (5.7 months) and now I expect ~150 day doubling times (5 months). Bluesky: Jeremy Diamond / @dmnd.me : There's a bur after you get past the toy/novelty uses where the learning curve gets steep fast. — I don't see a lot of discussion about how much more effort and scaffolding you need to put in to enable these products to run this long and actually produce good output metr.org/blog/2025-03... Ethan Mollick / @emollick : Since people are asking about the “exponential growth” bit - AI has a jagged frontier but on most hard careful measures of AI ability we are seeing rapid exponential gains until benchmarks get saturated. — Take METR long tasks: metr.org/blog/2025-03... [image] Alex Hern / @hern : METR has updated its LLM time-horizon chart for Claude 4.5 Sonnet and we're still above the trendline. Slowdown? What slowdown? metr.org/blog/2025-03... [image]
Context & Ripple Effects
METR's result adds a model-specific data point to a contested picture of coding-agent progress. Earlier coverage found that experienced developers using AI tools could be slower despite feeling faster, while a later assessment identified a “messiness tax” on lengthy real-world software tasks.
The result also arrives as Anthropic reported that its own employees use Claude heavily for debugging and code understanding, with self-reported productivity gains. That makes a longer task horizon relevant not just as a benchmark result but as a measure of how much work can plausibly be delegated before human review.
First-order effects
- Claude Opus 4.5 can be evaluated against substantially longer coding tasks at the 50% completion threshold than Claude Opus 4, strengthening Anthropic's case for agentic development workflows.
- Engineering teams gain a more concrete boundary for piloting longer-running Claude assignments, but the 50% metric still implies that verification and recovery remain integral to deployment.
Second-order effects
- Competing model providers and coding-tool vendors face pressure to demonstrate reliability on longer, end-to-end tasks rather than emphasize short benchmark wins.
- Buyers will increasingly evaluate AI coding tools by useful completed work and oversight burden; the earlier developer productivity study shows why perceived speed alone is an insufficient procurement metric.
Third-order effects
- If task horizons continue extending while reliability improves, software work may shift from tool-assisted coding toward workflows in which people scope, supervise, and validate multi-hour agent runs.
- The key competitive constraint would become dependable execution in messy production environments—not merely model capability—raising the value of evaluation, observability, and human-review infrastructure.
The trend: AI coding is moving toward longer-duration agent execution, with real-world reliability and cost per completed task becoming the decisive adoption measures.