/
Navigation
Chronicles
Browse all articles
Explore
Semantic exploration
Research
Entity momentum
Nexus
Correlations & relationships
Story Arc
Topic evolution
Drift Map
Semantic trajectory animation
Posts
Analysis & commentary
Pulse API
Tech news intelligence API
Browse
Entities
Companies, people, products, technologies
Domains
Browse by publication source
Handles
Browse by social media handle
Detection
Concept Search
Semantic similarity search
High Impact Stories
Top coverage by position
Sentiment Analysis
Positive/negative coverage
Anomaly Detection
Unusual coverage patterns
Analysis
Rivalry Report
Compare two entities head-to-head
Semantic Pivots
Narrative discontinuities
Crisis Response
Event recovery patterns
Connected
Search: /
Command: ⌘K
Embeddings: large
TEXXR

Chronicles

The story behind the story

days · browse · Enter similar · o open

OpenAI says it found widespread task issues in SWE-Bench Pro, estimates ~30% of tasks are broken, and retracts its earlier recommendation to adopt the benchmark

OpenAI

Context & Ripple Effects

OpenAI had previously promoted benchmark work beyond general-purpose evaluation, including a program for tailored domain benchmarks and a joint smart-contract-security benchmark with Paradigm. This reversal puts the reliability of the underlying task sets—not just model scores—at the center of that evaluation agenda.

Related coverage also shows OpenAI publicly revisiting failures in its own model-testing and release processes. The SWE-Bench Pro finding extends that pattern from model behavior to the quality controls behind the tests used to assess models.

First-order effects

  • OpenAI has withdrawn its earlier recommendation to adopt SWE-Bench Pro after identifying issues it says affect roughly 30% of its tasks, immediately weakening confidence in results derived from the benchmark.
  • Teams using SWE-Bench Pro to select, compare, or market coding agents now need to treat its scores as potentially contaminated until affected tasks are identified or replaced.

Second-order effects

  • Model providers and agent vendors that cite SWE-Bench Pro performance face pressure to re-run evaluations on cleaned task sets or substantiate claims with additional benchmarks.
  • The finding increases the value of narrower, domain-specific evaluations where task construction and validation can be more closely tied to real workflows, such as OpenAI's benchmark collaborations.

Third-order effects

  • If benchmark defects repeatedly alter model rankings, AI evaluation will shift away from single leaderboard scores toward benchmark portfolios, task audits, and disclosure of known limitations.
  • The episode may accelerate demands for more independent benchmark governance, since the credibility of model comparisons depends as much on dataset maintenance as on model capability.

The trend: AI evaluation is moving from broad leaderboard competition toward more rigorously validated, use-case-specific benchmarks and more transparent testing practices.

Discussion

  • @danielfein7 Daniel Fein on x
    SWE-Bench pro pronounced dead 2 hours before its poster presentation 😔 [image]
  • @minchoi Min Choi on x
    what [image]
  • @emollick Ethan Mollick on x
    The metrics discussion at OpenAI is a little confusing to me. I appreciate the clarification about bad benchmarks, but they spent a lot of money developing a very good benchmark of autonomous model ability at hard tasks, GDPval, and haven't reported it for GPT-5.6.
  • @himanshustwts Himanshu on x
    How to not make evals 101 [image]
  • @yacinemtb Kache on x
    lol
  • @theo @theo on x
    Agreed.
  • @matthewberman Matthew Berman on x
    This is pretty serious, right?
  • @openai @openai on x
    Our audit of SWE-Bench Pro found that a meaningful share of public tasks contain issues that can distort results. Some correct solutions fail because of hidden requirements, contradictory instructions, overly strict tests, or incomplete grading criteria. [image]
  • @openai @openai on x
    We audited SWE-Bench Pro, one of the most widely used AI coding benchmarks, and found it no longer reliably measures frontier coding capability. We find 30% of SWE-Bench Pro tasks to be broken, and are retracting our previous recommendation that the research community use it as
  • @openai @openai on x
    To audit SWE-Bench Pro, we used model-based investigator agents alongside independent reviews from five independent experienced software engineers. That helped us examine tasks at scale while keeping expert judgment at the center. [image]
  • @_amankishore Aman on x
    Oh that benchmark you were beating us on? Actually it's invalid [image]
  • @stevehou Steve Hou on x
    Interesting. The audited are auditing the auditors.
  • @matei_zaharia Matei Zaharia on x
    One reason we built custom coding and data agent benchmarks internally at Databricks (e.g. https://www.databricks.com/...). Academic benchmarks are great and people will build better ones, but you also care about YOUR tasks, which are often different. Each company needs its own “…
  • @deepdishenjoyer @deepdishenjoyer on x
    lmaooooooo i guess that's one approach to fable kicking your ass
  • Bilal Aslam Bilal Aslam on linkedin
    Super interesting read on the performance of frontier models on coding tasks — as measured against Databricks' code base  —  https://lnkd.in/...
  • r/singularity r on reddit
    OpenAI finds ~30% of tasks in SWE Bench Pro are broken
  • r/mlscaling r on reddit
    SWE-Bench Pro is now saturated at 70%