Clearview CEO says it has collected over 10B images from across the web and is building ways for police to find people, including deblurring and mask removal
In an interview with WIRED, CEO Hoan Ton-That said the company has scraped 10 billion photos from the web—and developed new ways to aid police surveillance. Tweets: @zittrain , @jetjocko , @brbarrett , @iwillleavenow , @onekade , and @willknight Tweets: Jonathan Zittrain / @zittrain : The public and private failures to confront the scraping of 10 billion photos of people, with tags — to facilitate instant facial recognition of anyone for the rest of their lives — will rank as one of the biggest unforced and irreversible losses of privacy in thirty years. https://twitter.com/... Adam Rogers / @jetjocko : Strong tautilogitronics vibes from “we scraped 10 billion selfies for surveillance” company Clearview and my colleague @willknight https://www.wired.com/... https://twitter.com/... Brian Barrett / @brbarrett : Clearview AI now has more than 10 billion images at its disposal, and is introducing “deblur” and “antimasking” tools. What could go right? https://www.wired.com/... Techni-Calli / @iwillleavenow : You think someday an oversight body is going to do literally anything about Clearview AI besides the occasional strongly-worded letter or...? https://www.wired.com/... @onekade : An image not good enough to run through facial recognition? No worries! Just make shit up! https://www.wired.com/... https://twitter.com/... Will Knight / @willknight : My latest story: The man behind ClearviewAI, a controversial face recognition tool that lets police find people online, says he's developing new AI features to make it more powerful. Some say it could also make it even more problematic. https://www.wired.com/...
Context & Ripple Effects
Clearview has been on a straight scaling line since the original [[a:949805|New York Times exposé put its 3-billion-image scrape and 600-agency customer list on the record in January 2020]]. A leaked investor deck a year ago showed the database already past 10 billion and management targeting 100 billion photos within a year. Today's interview with CEO Hoan Ton-That confirms the interim number and adds a capability layer: deblurring and mask removal built specifically so police can identify people who were previously unmatchable in low-quality or covered footage.
First-order effects
- Police users gain identification reach into exactly the imagery that facial recognition historically failed on — blurred CCTV and masked faces — while every person with photos online becomes searchable against a corpus more than three times the size reported at launch.
- Ton-That is publicly committing to the growth trajectory rather than defending against it, converting what was investigative reporting into confirmed company roadmap.
Second-order effects
- Platforms whose content feeds the scrape — Facebook, YouTube and similar sites named in earlier coverage — face renewed pressure to technically block or legally contest bulk image harvesting, since their users' photos are now inputs to police tooling without consent.
- Rival surveillance vendors must match both the dataset scale and the preprocessing tricks (deblurring, antimasking) or concede accuracy leadership on the hardest real-world footage.
Third-order effects
- If the pattern holds, biometric identification becomes permanent and retroactive: as Jonathan Zittrain's reaction notes, a scrape of tagged photos enables instant recognition of anyone for the rest of their lives — an irreversible privacy loss that regulators will have to address through the unresolved question of whether publicly posted data implies consent to mass facial indexing.
- The gap between what platforms permit and what scrapers can build accelerates the push for an explicit legal boundary on training and search databases built from public data.
The trend: Facial recognition is scaling from curated databases toward total web-scale indexes augmented by image-restoration AI, with law enforcement demand outpacing any consent framework governing the source data.