OpenAI says it found widespread task issues in SWE-Bench Pro, estimates ~30% of tasks are broken, and retracts its earlier recommendation to adopt the benchmark
OpenAI
Context & Ripple Effects
OpenAI had previously promoted benchmark work beyond general-purpose evaluation, including a program for tailored domain benchmarks and a joint smart-contract-security benchmark with Paradigm. This reversal puts the reliability of the underlying task sets—not just model scores—at the center of that evaluation agenda.
Related coverage also shows OpenAI publicly revisiting failures in its own model-testing and release processes. The SWE-Bench Pro finding extends that pattern from model behavior to the quality controls behind the tests used to assess models.
First-order effects
OpenAI has withdrawn its earlier recommendation to adopt SWE-Bench Pro after identifying issues it says affect roughly 30% of its tasks, immediately weakening confidence in results derived from the benchmark.
Teams using SWE-Bench Pro to select, compare, or market coding agents now need to treat its scores as potentially contaminated until affected tasks are identified or replaced.
Second-order effects
Model providers and agent vendors that cite SWE-Bench Pro performance face pressure to re-run evaluations on cleaned task sets or substantiate claims with additional benchmarks.
The finding increases the value of narrower, domain-specific evaluations where task construction and validation can be more closely tied to real workflows, such as OpenAI's benchmark collaborations.
Third-order effects
If benchmark defects repeatedly alter model rankings, AI evaluation will shift away from single leaderboard scores toward benchmark portfolios, task audits, and disclosure of known limitations.
The episode may accelerate demands for more independent benchmark governance, since the credibility of model comparisons depends as much on dataset maintenance as on model capability.
The trend: AI evaluation is moving from broad leaderboard competition toward more rigorously validated, use-case-specific benchmarks and more transparent testing practices.
The metrics discussion at OpenAI is a little confusing to me. I appreciate the clarification about bad benchmarks, but they spent a lot of money developing a very good benchmark of autonomous model ability at hard tasks, GDPval, and haven't reported it for GPT-5.6.
Our audit of SWE-Bench Pro found that a meaningful share of public tasks contain issues that can distort results. Some correct solutions fail because of hidden requirements, contradictory instructions, overly strict tests, or incomplete grading criteria. [image]
We audited SWE-Bench Pro, one of the most widely used AI coding benchmarks, and found it no longer reliably measures frontier coding capability. We find 30% of SWE-Bench Pro tasks to be broken, and are retracting our previous recommendation that the research community use it as
To audit SWE-Bench Pro, we used model-based investigator agents alongside independent reviews from five independent experienced software engineers. That helped us examine tasks at scale while keeping expert judgment at the center. [image]
One reason we built custom coding and data agent benchmarks internally at Databricks (e.g. https://www.databricks.com/...). Academic benchmarks are great and people will build better ones, but you also care about YOUR tasks, which are often different. Each company needs its own “…