/
Navigation
Chronicles
Browse all articles
Explore
Semantic exploration
Research
Entity momentum
Nexus
Correlations & relationships
Story Arc
Topic evolution
Drift Map
Semantic trajectory animation
Posts
Analysis & commentary
Pulse API
Tech news intelligence API
Browse
Entities
Companies, people, products, technologies
Domains
Browse by publication source
Handles
Browse by social media handle
Detection
Concept Search
Semantic similarity search
High Impact Stories
Top coverage by position
Sentiment Analysis
Positive/negative coverage
Anomaly Detection
Unusual coverage patterns
Analysis
Rivalry Report
Compare two entities head-to-head
Semantic Pivots
Narrative discontinuities
Crisis Response
Event recovery patterns
Connected
Search: /
Command: ⌘K
Embeddings: large
TEXXR

Chronicles

The story behind the story

days · browse · Enter similar · o open

Microsoft has pulled its facial recognition database, MS Celeb, which contained 10M+ images of ~100K individuals scraped under the Creative Commons license

Madhumita Murgia / Financial Times :

Financial Times Madhumita Murgia

Context & Ripple Effects

The pull is the sharpest move yet in a story the Financial Times has been building all spring: an [[a:940864|April survey of the face datasets available on request from universities and the US government]] showed how much of commercial facial recognition rests on images collected for other purposes. MS Celeb was the biggest example — 10M+ images of ~100K people swept up under Creative Commons licenses that were never meant to authorize biometric reuse.

First-order effects

  • Researchers and companies that trained models on MS Celeb lose their primary training corpus overnight, and Microsoft removes its own brand from the most-cited example of license-scraped biometric data.
  • The move lands weeks after Microsoft publicly refused a California law enforcement request over bias concerns while still selling the tech to a US prison — so the company is retreating from the data layer while keeping the product layer.

Second-order effects

  • Every university and government lab holding similar request-based face datasets now faces the same question Microsoft just answered for itself: whether a Creative Commons tag constitutes consent, and whether their holdings are a reputational liability.
  • Creative Commons itself is pushed into clarifying that its licenses govern attribution and redistribution of works, not the harvesting of identifiable faces — a boundary the licensing framework was never designed to draw.

Third-order effects

  • If the pattern holds, the face-recognition data supply chain gets rebuilt around explicit collection rather than scraping: Microsoft went on to strip gender, age, and emotion inference from Azure under its Responsible AI Standard (the 2022 removal), and Meta deleted over a billion face scans outright — suggesting dataset withdrawal is the first step in a broader corporate unwind of facial analysis.
  • For regulators, each voluntary pull becomes evidence that the permission gap between copyright law and biometric privacy cannot be left to corporate discretion, strengthening the case for rules that treat a face as something no existing license can grant.

The trend: The companies that built facial recognition on repurposed public images are dismantling that data supply chain themselves, converting license-scraped datasets from an asset into a liability.