Missing data hindering replication of AI studies as survey of 400 algorithms presented at major conferences finds just 6% included code, 30% included test data
Last year, computer scientists at the University of Montreal (U of M) in Canada were eager to show off a new speech recognition algorithm … Tweets: @hanno , @jorendorff , @silverjacket , @meaningness , and @mattmight Tweets: @hanno : This detail of the recent article about AI+replication tells you one thing: Just trying to replicate isn't enough, replications need to be preregistered and published regardless of outcome. http://www.sciencemag.org/... http://twitter.com/... Every_jorendorff / @jorendorff : “Researchers say there are many reasons for the missing details: The code might be a work in progress, owned by a company, or held tightly by a researcher eager to stay ahead of the competition.” you're all fired http://twitter.com/... Matthew Hutson / @silverjacket : Artificial intelligence faces reproducibility crisis. My story in this week's @sciencemagazine: http://www.sciencemag.org/... David Chapman / @meaningness : AI doesn't replicate. Having worked in the field, I can usually see why a paper's result is nonsense, but the public can't, and many researchers can't. http://twitter.com/... Matt Might / @mattmight : This looks like a job for the CRAPL: http://matt.might.net/... Papers which don't release code and data shouldn't be published. No exceptions. End of story. This is a huge embarrassment for the field of computer science. http://twitter.com/...
Context & Ripple Effects
This Science survey of 400 algorithms from major conferences — just 6% with code and 30% with test data — was the first hard measurement of what had been an anecdotal complaint about AI research. The University of Montreal speech-recognition example it opens with shows the problem hitting even top-lab work: without code or data, 'trying to replicate' produces nothing publishable.
The finding aged into a running theme rather than a one-off: researchers formally acknowledged a reproducibility crisis by late 2019 (Wired's coverage), and by 2020 the critique widened to unequal access to code, proprietary data, and hardware (MIT Technology Review) — with peer reviewers left carrying ethical oversight too (New Yorker, 2021).
First-order effects
- Authors of the surveyed conference papers face immediate credibility pressure: with only 6% releasing code, the default assumption shifts toward 'unverifiable' for any result published without artifacts.
Second-order effects
- Conference organizers and reviewers are pushed to make code and test-data release a review criterion, since the survey gives them a baseline number to enforce against; labs hoarding proprietary datasets gain a competitive moat over academic groups who cannot match closed resources.
Third-order effects
- If artifact release becomes standard, verification infrastructure — preregistration, published replications regardless of outcome, shared benchmarks — turns into shared field infrastructure, and the gap between well-resourced industry labs and academia widens along access-to-data lines rather than talent lines.
The trend: AI research is moving from publication-as-claim toward publication-as-artifact, with reproducibility pressure accumulating across surveys, crisis acknowledgments, and transparency critiques since 2018.