METR study: experienced open-source developers using Cursor, Claude, and other AI tools were 19% slower to complete tasks, despite thinking they were 20% faster
Study Shows That Even Experienced Developers Dramatically Overestimate Gains — The buzz about AI coding tools is unrelenting.
Context & Ripple Effects
This result challenges productivity claims built largely on developer perception. In later coverage, Anthropic employees reported a 50% productivity boost from Claude use, concentrated in debugging and code understanding—an instructive contrast with this study’s measured task times self-reported Claude productivity gains.
The discrepancy also fits subsequent evidence that AI-tool use can be weakest where verification matters: an Anthropic experiment found its largest performance decline in debugging tasks debugging showed the largest performance decline. The key issue is not whether developers use these tools, but whether apparent speed survives task-level measurement.
First-order effects
- Experienced open-source developers in the study took 19% longer with Cursor, Claude, and similar tools, while believing they were 20% faster—making subjective productivity reports unreliable for these tasks.
- Teams deploying AI coding assistants need to distinguish perceived fluency from completed, validated work before treating tool adoption as a capacity gain.
Second-order effects
- Tool vendors and engineering leaders face pressure to demonstrate outcomes on real repositories and workflows, rather than relying on adoption, satisfaction, or self-reported time savings.
- If review, debugging, or correction absorbs the lost time, the relevant purchasing metric shifts toward quality- and security-adjusted output, not code-generation speed alone.
Third-order effects
- The AI coding market may increasingly compete on measurable task completion and verification, with tools that reduce downstream debugging carrying more value than tools that merely accelerate first drafts.
- The study points to a durable measurement gap: organizations that instrument end-to-end engineering work may make more disciplined AI spending decisions than those that extrapolate from developer sentiment.
The trend: AI coding assistants are moving from a novelty-driven adoption cycle toward scrutiny of their net, validated productivity on real software work.