Study: AI that detects eye diseases is mostly trained on patients in US, Europe, and China, which can make it ineffective for other racial groups and minorities
Will Knight / Wired :
Context & Ripple Effects
The training-data gap in eye-disease AI lands a month after CMS began paying doctors to use AI that diagnoses diabetic eye disease, meaning US reimbursement infrastructure is scaling up for exactly the kind of tool this study says was built on narrow patient populations.
The finding extends a documented pattern: researchers had already shown an algorithm used across the US healthcare system systematically shortchanged black patients (the widely used needs-assessment algorithm study), and Google's diabetic retinopathy screener struggled in real-world trials in Thailand despite strong lab accuracy — same disease area, same gap between benchmark performance and field performance for non-representative populations.
First-order effects
- Developers of retinal screening tools now face validation questions before deploying to populations outside the US, Europe, and China, since models tuned to those cohorts may miss disease presentations in other racial groups.
- Health systems in underrepresented regions that adopt these off-the-shelf tools risk getting screenings that look accurate on paper but underperform for their own patients.
Second-order effects
- Payer programs like CMS's new eye-disease AI reimbursement create financial incentives to deploy broadly, so a biased model doesn't just sit unused — it scales through billing channels, forcing buyers to demand demographic performance data vendors haven't historically provided.
- Vendors competing for clinical deployments will need region- and population-specific datasets, raising acquisition costs relative to competitors who reuse US/European/Chinese data.
Third-order effects
- If the pattern holds — from the needs-assessment algorithm through retinal screening to LLM symptom assessment — regulatory approval for medical AI will likely shift from aggregate accuracy toward demonstrated per-population performance, restructuring who can afford to enter clinical markets.
- The broader constraint identified by Wired's own coverage — medical data being scarcer and more complex than web data — points toward medical AI consolidating among players who can fund population-level data collection rather than those with the best algorithms.
The trend: Medical AI is moving from benchmark-driven claims toward population-specific validation, as repeated bias findings force regulators and payers to treat training-data representativeness as a safety requirement.