/
Navigation
Chronicles
Browse all articles
Explore
Semantic exploration
Research
Entity momentum
Nexus
Correlations & relationships
Story Arc
Topic evolution
Drift Map
Semantic trajectory animation
Posts
Analysis & commentary
Pulse API
Tech news intelligence API
Browse
Entities
Companies, people, products, technologies
Domains
Browse by publication source
Handles
Browse by social media handle
Detection
Concept Search
Semantic similarity search
High Impact Stories
Top coverage by position
Sentiment Analysis
Positive/negative coverage
Anomaly Detection
Unusual coverage patterns
Analysis
Rivalry Report
Compare two entities head-to-head
Semantic Pivots
Narrative discontinuities
Crisis Response
Event recovery patterns
Connected
Search: /
Command: ⌘K
Embeddings: large
TEXXR

Chronicles

The story behind the story

days · browse · Enter similar · o open

Researchers detail health care AI's reproducibility issues; a 2021 review of 500+ papers: health ML models perform especially poorly on reproducibility measures

Nature Emily Sohn

Context & Ripple Effects

This finding lands at the end of a five-year arc. A 2018 survey of 400 algorithms presented at major conferences found just 6% included code and 30% test data; by 2019 researchers were openly acknowledging a reproducibility crisis, and 2020 criticism widened to unequal access to proprietary code, data, and hardware. What is new here is scope: a 2021 review of 500+ papers shows health-related ML performing especially poorly on reproducibility measures — the domain where model errors carry patient consequences.

First-order effects

  • Hospital and clinic teams evaluating published health ML models cannot assume reported results will hold on their own populations, since the underlying code and data behind most studies are not available to check.
  • Researchers publishing health AI now face direct scrutiny over whether their artifacts — code, test data, preprocessing steps — can be released at all, given how many prior studies could not be replicated.

Second-order effects

  • Vendors selling clinical AI into health systems inherit the credibility gap: buyers who cannot independently reproduce academic claims will lean harder on vendor-supplied validation, shifting due-diligence burden onto procurement.
  • Journals and funders covering health ML come under pressure to mandate artifact release, because the alternative — trusting unreplicable results in medicine — is harder to defend than in general AI research.

Third-order effects

The trend: AI's reproducibility crisis, first documented in general research, is hardening into a formal validation requirement where it matters most — clinical deployment.

Discussion

  • @naturecareers @naturecareers on x
    “Given the exploding nature and how widely these things are being used, I think we need to get better more quickly than we are,” says @GreeneScientist #AI https://go.nature.com/3ZrGZJC
  • @rielymd @rielymd on x
    Excellent article highlighting issues around AI in health care. Notes a number of challenges: “a major issue is the relative scarcity of publicly available data sets in medicine.” Reproducibility issues that haunt health-care AI https://www.nature.com/...
  • @garymarcus Gary Marcus on x
    “The reproducibility issues that haunt health-care AI” / ⁦and let me recommend @jpineau1⁩'s reproducibility checklist yet again! https://www.nature.com/...
  • @pkedrosky Paul Kedrosky on x
    Some of the best-performing AI algorithms for interpreting CT scans are, when re-tested, no better than coin flips. The reproducibility issues that haunt health-care AI https://www.nature.com/... #xp
  • @tedescosalvo Salvatore Tedesco on x
    “Health-related ML models perform particularly poorly on reproducibility measures relative to other ML disciplines...a major issue is the relative scarcity of publicly available datasets in medicine with the result that biases/inequities become entrenched” https://www.nature.com/…
  • @natureportfolio @natureportfolio on x
    .@Nature reports on the move towards increased reproducibility in health-care AI, including strategies such as greater algorithmic transparency and promoting checklists to avoid common errors. https://go.nature.com/3XpEnKz
  • @bermaninstitute @bermaninstitute on x
    The reproducibility issues that haunt health-care AI: Health-care systems are rolling out artificial-intelligence tools for diagnosis and monitoring. But how reliable are the models? https://www.nature.com/...
  • @realhayman Hayman Buwaneswaran Buwan on x
    The reproducibility issues that haunt healthcare #AI - Healthcare systems are rolling out artificial intelligence tools for diagnosis and monitoring. But how reliable are the models? https://www.nature.com/... #digitalhealth #medtech #ML @medtechshow via @clemoscatarina
  • @broadhurstdavid David Broadhurst on x
    Great article on a very important subject. Unfortunately you could also change “health-care AI” to “'omics AI” with similar observations. #reproducibility #omics #metabolomics #proteomics #genomics #AI #MachineLearning https://www.nature.com/...
  • @cmichaelgibson C. Michael Gibson MD on x
    The reproducibility issues that haunt health-care AI: When algorithms with 90% “accuracy” are tested on fresh new sets of data, the accuracy often drops to 60-70% (a little better than a coin toss) https://www.nature.com/...
  • @vickyhellon Vicky Hellon on x
    Great article on reproducibility issues with AI in healthcare 🤖🩺 Potential solutions include making models and data publicly available, eliminating redundancy between training and testing datasets and implementing checklists e.g from the @EQUATORNetwork https://www.nature.com/...