Research: machine learning techniques today are designed to find patterns and always make a prediction, making a lot of ML-based scientific research unreliable
Artificial intelligence is being applied with undue haste to analyse data in some areas of biomedical research …
Context & Ripple Effects
This FT-reported finding lands mid-arc in a running critique of applied AI. Back in 2017, MIT Technology Review flagged that [[a:918139|deep learning systems are neither understandable to their creators nor accountable to their users]] — a transparency problem. This research adds a structural one: because ML techniques are built to find patterns and always emit a prediction, they manufacture confident answers even where no reliable signal exists, which is precisely the wrong default for biomedical data analysis.
The concern compounds rather than stands alone. Weeks after this report, researchers argued that deep learning is nearing its limits and needs new approaches for further progress; by 2022, Wired reported that AI in health care had underdelivered because medical data is more complex and scarcer than web data, producing misleading results; and in 2024 researchers warned that training AI on synthetic data could drive model degradation over time. Together these sketch a field discovering that its core machinery fails quietly in high-stakes, data-poor settings.
First-order effects
- Biomedical researchers using off-the-shelf ML tools face unreliable findings by construction — the tools' always-predict design means spurious patterns get published as discoveries, putting replication and credibility of affected studies at risk.
- Journals, funders, and labs applying AI to biomedical data must now treat every ML-generated result as unvalidated until independently confirmed, raising the cost of the analysis pipelines many had adopted at speed.
Second-order effects
- Vendors selling ML analysis into health care and life sciences face a trust deficit on top of the existing problem that medical data's complexity and scarcity already produce misleading results — buyers will demand evidence the tool knows when not to predict.
- The pressure pushes method development toward approaches that can abstain or quantify uncertainty, giving traction to the argument that deep learning itself is nearing its limits and new architectures are needed.
Third-order effects
- If the pattern holds — forced predictions in science, degraded models trained on synthetic data, opaque systems accountable to no one — AI adoption in research and medicine shifts from rapid deployment toward validation-first regimes, with reliability becoming the gating criterion rather than raw predictive power.
- A sustained reliability gap between web-scale AI successes and scientific applications would bifurcate the field: consumer-facing AI scales on abundant data while scientific AI stalls until methods are rebuilt for scarce, complex data.
The trend: Across biomedicine, health care, and model training itself, AI is hitting a reliability reckoning in which the discipline of knowing when a model should refuse to answer becomes as important as its ability to predict.