A comprehensive look at the plethora of face datasets, available upon request from universities and US government, used to train facial recognition software
Madhumita Murgia / Financial Times : Tweets: @matt_odell , @_jpfox , @arcalinea , @anncavoukian , @hoofnagle , @hoofnagle , and @eff Tweets: Matt Odell / @matt_odell : “Now somebody's face is used as a tracking number to watch them as they move across locations on video, which is a huge shift...But no one ever stopped to think if it was ethical to collect images of people's weddings & family photo albums with children” http://www.ft.com/... JP Fox / @_jpfox : ‘...a plethora of face repositories have sprung up, containing images manually culled and bound together from sources as varied as university campuses, town squares, markets, cafés, mugshots and social-media sites such as Flickr, Instagram and YouTube.’ https://www.ft.com/... Jay Graber / @arcalinea : Unfortunately for those who wish to opt out of surveillance, you can't encrypt your face. http://ft.com/... Ann Cavoukian, Ph.D. / @anncavoukian : Unfortunately, it's not that simple because facial recognition is so pervasive and often invisible: you may never know it's taking place, giving you no opportunity to challenge it. Facial recognition may be the most sensitive biometric data out there next to your DNA. http://twitter.com/... Chris Hoofnagle / @hoofnagle : This article discusses how EFF's Jillian York is in an IARPA-funded face recognition database, with sources including youtube stills and photos available from Google going back to 2008 http://www.ft.com/... Chris Hoofnagle / @hoofnagle : It's 2019 and we're just now realizing that face recognition projects are hoovering up images from all “public” (i.e. online) sources. Here's a hint: if a company has a picture of you, it is running face recognition on it. Why wouldn't they? http://www.ft.com/... @eff : Your photo could be in a face recognition database without you knowing. Last month, EFF's @jilliancyork discovered her photos, along with other EFF staff, journalists, and activists, were being used to train algorithms to recognize “suspects of interest.” http://www.ft.com/...
Context & Ripple Effects
Before Clearview AI made scraping famous, the raw material for facial recognition was already circulating through quieter channels: face repositories assembled from Flickr, Instagram, YouTube and mugshot databases and handed out by universities and US government agencies on request. The FT's mapping shows how little consent sat underneath the field — EFF's Jillian York found photos of her own staff, journalists and activists inside training sets for suspect-recognition algorithms.
The piece lands between two markers in the arc: NIST had already been caught benchmarking systems on images of immigrants, visa applicants, abused children and dead people without consent (Slate's March 2019 report), and months later Clearview would show what happens when the same logic scales to billions of scraped images. Privacy figures like Ann Cavoukian frame the core question — nobody stopped to ask whether collecting wedding albums and children's photos was ethical.
First-order effects
- Universities and US government agencies are directly exposed: they are the distributors of these on-request datasets, and EFF's discovery that its own staff appear in them turns academic infrastructure into a privacy incident.
- The photographed subjects — wedding guests, family members, children pulled from social media — have no notification or removal path, since the datasets were compiled without their knowledge.
Second-order effects
- Platforms whose content feeds these repositories — Flickr, Instagram, YouTube — face pressure to treat bulk image harvesting as a violation rather than fair use, a line Clearview's later scraping made impossible to ignore.
- Law enforcement and intelligence buyers of trained models inherit tainted provenance: if the training data was gathered unethically, every downstream suspect-recognition deployment carries that liability.
Third-order effects
- If the pattern holds, the field splits along a consent boundary: datasets built on explicit permission versus those scraped from public posts, with regulators eventually forced to decide whether 'publicly posted' means 'usable for biometric training'.
- Government agencies funding and distributing these collections become the de facto standards-setters for biometric AI, meaning ethics debates move from companies to the state bodies supplying the data.
The trend: Facial recognition is outgrowing its curated-dataset origins toward mass harvesting of public photos, forcing a reckoning over whether anything posted online is fair training material.