Research: AI's ability to complete lengthy software engineering tasks has doubled roughly every six months, but there is a “messiness tax” for real-world tasks
METR has had a very influential work by Kwa and West et al on measuring AI's ability to complete long tasks. X: @kirillzzy , @boazbaraktcs , @benshindel , @jasonfurman , @jasonfurman , and @sama X: Kirill Avery / @kirillzzy : if the bots are going to be doing human jobs, then the humans better be getting paid some other way 🤔 Boaz Barak / @boazbaraktcs : Wrote a blog post with my thoughts on how AI could impact economic growth. [image] Ben / @benshindel : I realize this view of history is extremely common these days, but I think ppl underestimate the technological and cultural progress between 100k ya and the rise of cities 5k ya, and again from there to the height of the Bronze Age, and then to the classical era, and so forth. [image] Jason Furman / @jasonfurman : I'm thinking of the engineering studies about ways that households could weatherize their homes, reduce electricity bills & help climate. Then economists study actual people and they did everything so imperfectly the actual results were much more disappointing. Is AI like that? Jason Furman / @jasonfurman : Fascinating on the economic implications of AI from @boazbaraktcs. Boaz is more optimistic than me. But also vastly more knowledgable about AI than me. One issue I wonder about is the engineer (studying how it works in a lab) vs. the economist (studying people in the wild). [image] Sam Altman / @sama : interesting post from @boazbaraktcs: https://windowsontheory.org/ ...
Context & Ripple Effects
METR’s work on long-horizon software tasks sits alongside a prior finding that experienced open-source developers were slower with AI coding tools despite expecting a speedup. The new result separates rapid gains on task completion from the frictions of deploying those capabilities in real engineering environments.
The “messiness tax” also gives substance to a broader debate over whether model progress transfers cleanly to on-the-job learning and generalization, rather than treating benchmark improvement as a direct measure of workplace substitution.
First-order effects
- AI labs and coding-tool vendors gain evidence of fast improvement on lengthy software tasks, while the reported real-world penalty limits how directly that progress can be translated into production-work claims.
- Engineering teams evaluating AI agents must distinguish controlled task performance from work involving the incomplete context, coordination, and irregular inputs implied by the messiness tax.
Second-order effects
- Vendors will be pressured to demonstrate results on realistic repositories and workflows, not only on cleaner long-task evaluations; procurement comparisons will increasingly turn on reliability in context.
- The gap between capability and deployment value reinforces demand for human review and workflow design, echoing the role of human subject-matter work behind AI systems rather than making that work disappear immediately.
Third-order effects
- If long-task capability continues improving at this pace while real-world transfer lags, software automation will likely be adopted unevenly: first where tasks and environments can be made legible to agents, later in more ambiguous work.
- The key competitive metric may shift from headline model capability toward cost per completed, verified task—linking AI progress to the economics of deployment rather than benchmark scores alone.
The trend: AI agents are advancing rapidly on extended tasks, but the economic payoff is increasingly determined by how well models handle the unstructured conditions of actual work.