/
Navigation
Chronicles
Browse all articles
Explore
Semantic exploration
Research
Entity momentum
Nexus
Correlations & relationships
Story Arc
Topic evolution
Drift Map
Semantic trajectory animation
Posts
Analysis & commentary
Pulse API
Tech news intelligence API
Browse
Entities
Companies, people, products, technologies
Domains
Browse by publication source
Handles
Browse by social media handle
Detection
Concept Search
Semantic similarity search
High Impact Stories
Top coverage by position
Sentiment Analysis
Positive/negative coverage
Anomaly Detection
Unusual coverage patterns
Analysis
Rivalry Report
Compare two entities head-to-head
Semantic Pivots
Narrative discontinuities
Crisis Response
Event recovery patterns
Connected
Search: /
Command: ⌘K
Embeddings: large
TEXXR

Chronicles

The story behind the story

days · browse · Enter similar · o open

Google's AI screening tool for diabetic retinopathy, trialed in Thailand, proved impractical in real-life testing, despite high theoretical accuracy

AI is frequently cited as a miracle workers in medicine, especially in screening processes, where machine learning models boast expert-level skills in detecting problems.

TechCrunch Devin Coldewey

Context & Ripple Effects

Google has spent years building toward autonomous eye screening: DeepMind's five-year project on a million anonymous NHS eye scans supplied the research base, the FDA cleared the first retinal-photo diagnostic needing no doctor interpretation in 2018, and Google then took the tool into the field with an India screening program for at-risk diabetics.

The Thailand trial is the reality check in that arc: a model with expert-level theoretical accuracy ran into practical obstacles once deployed in actual clinics. That matters because the same gap between benchmark scores and bedside usability now sits under every medical-AI deployment Google is scaling, from retina screening to the mammography results it publicized earlier in 2020.

First-order effects

  • Google's field deployments — the India screening program and the Thai trial sites — stall at the last mile: nurses and clinic staff bear the workflow cost of operating a tool whose accuracy was proven only under controlled conditions.
  • The mammography study's headline accuracy claims now carry a credibility discount, since the same lab-to-clinic gap has been demonstrated inside Google's own portfolio.

Second-order effects

  • Reimbursement pulls against practicality: CMS's move to pay doctors for AI-based eye-disease and stroke diagnosis creates a financial incentive to adopt tools that the Thai trial suggests may not survive contact with real workflows, forcing payers to weigh usability alongside approval status.
  • Rival diagnostic-AI vendors gain an argument for selling deployment support and integration services rather than raw model accuracy, reframing competition around operational fit.

Third-order effects

  • If the pattern holds, validation in medical AI shifts from retrospective accuracy studies to prospective operational evidence — meaning regulators like the FDA may need approval criteria that test workflow feasibility, not just detection rates.
  • Screening programs in resource-constrained health systems become the proving ground that determines whether autonomous diagnostics scale globally or remain confined to well-resourced clinics.

The trend: Medical AI is crossing from the accuracy-benchmark era into a deployment era, where workflow practicality in real clinics — not model precision — becomes the binding constraint on adoption.

Discussion

  • @rakeshlobster Rakesh Agrawal on x
    Product management is about marrying technology with understanding psychology. This is a great example. Put abstract tech out in real world and things fall apart. https://techcrunch.com/...
  • @maxalittle Max Little on x
    Even in the narrowest of medical decision problems (detecting diabetic retinopathy), Google's “AI” was unable to function on 20% of real-world images, causing increased work and unnecessary delays for both staff and patients. @GaryMarcus https://www.technologyreview.com/ ...
  • @gquaggiotto Giulio Quaggiotto on x
    “An AI system needs to fit into a process where sources of uncertainty are discussed rather than simply rejected” https://www.technologyreview.com/ ... Reality check from Thailand via @CoHuBiCoL1 @cabitzaf
  • @cabitzaf Federico Cabitza on x
    “Of course, we don't want an AI to make a bad call. But human doctors disagree all the time — and that's fine. An AI system needs to fit into a process where sources of uncertainty are discussed rather than simply rejected.” https://www.technologyreview.com/ ...
  • @melmitchell1 Melanie Mitchell on x
    Good reality check. Accuracy on a benchmark dataset doesn't necessarily reflect real-world complexities. https://twitter.com/...