Source: NIH and Google planned to launch a dataset of 100K chest X-rays in 2017, scrapped project days before launch, after finding personal info on some images
Fumbled project with NIH highlights potential pitfalls of Google's ambitions with sensitive health data
Context & Ripple Effects
The Washington Post's reporting reveals that Google and the [[a:|NIH]] scrapped a planned public release of 100,000 chest X-rays at the last minute in 2017 after discovering personal information embedded on some images — an early stumble that predates the company's bigger health-data pushes. It landed just as Google's health ambitions were surfacing publicly: days later, sources described the partnership with Ascension to analyze records of millions of patients under Project Nightingale.
The near-launch also sits inside a documented pattern of medical imaging leaking at scale — researchers had already found 16M identifiable scans online, and later reporting put unsecured medical images at over a billion — meaning the de-identification problem Google caught is one hospitals themselves routinely fail to catch.
First-order effects
- Researchers who were set to use the 100K-image dataset lost access to it entirely, and NIH and Google absorbed the reputational cost of a privacy failure caught internally days before going public.
Second-order effects
- The episode feeds the scrutiny around Project Nightingale and the Ascension deal, handing critics a concrete example that Google's internal review processes are the only barrier between patient data and public exposure.
Third-order effects
- If large-scale medical AI datasets keep hitting de-identification failures, curation shifts from an academic task to a compliance function — favoring players who can absorb audit costs and pushing public-private dataset releases toward stricter consent frameworks rather than scraped imaging archives.
The trend: Medical imaging is becoming AI's most contested data frontier, with de-identification failures at both hospitals and research institutions forcing every large dataset effort into a de facto permission fight.