NIST's benchmark test for facial recognition systems uses images of immigrants, US visa applicants, abused children, and dead people, without consent
If you thought IBM using “quietly scraped” Flickr images to train facial recognition systems was bad, it gets worse. Tweets: @rajit_h and @profwernimont Tweets: Rajit Hewagama / @rajit_h : The very group the U.S. government has tasked with regulating the facial recognition industry is perhaps the worst offender when it comes to using images sourced without the knowledge of the people in the photographs. http://slate.com/... Jacqueline Wernimont / @profwernimont : I'm fascinated by the number of “what's the harm, it's just being used as data” responses to this piece on the non-consensual use of child exploitation, law enforcement encounter, and visa application pictures to test #AI http://slate.com/... w/@drnikki @farbandish
Context & Ripple Effects
The Slate report lands a week after coverage of [[a:939436|IBM and other companies training facial recognition on scraped Creative Commons Flickr photos]] made corporate data practices the story of the month. What makes this installment different is the subject: NIST runs the [[a:939395|Facial Recognition Vendor Test program that shapes purchasing decisions across US agencies and businesses]], so the body tasked with judging the industry is itself sourcing images — of immigrants, visa applicants, abused children, and deceased people — without consent.
The finding also feeds directly into the broader dataset question the Financial Times took up a month later, cataloguing the face datasets available on request from universities and the US government. If the government's own referee operates outside any consent norm, the de facto standard for everyone else is already set.
First-order effects
- NIST's position as arbiter of facial recognition accuracy is undercut at the moment its FRVT rankings carry the most commercial weight — vendors ranked by the benchmark are effectively competing on tests built from non-consensual images.
- The people in the datasets — immigrants, visa applicants, abused children, and the dead — have no mechanism to know they were used or to opt out, since the images were sourced through government-held collections rather than public web scraping.
Second-order effects
- Companies named in the adjacent consent controversies, IBM most prominently, gain an uncomfortable defense-in-kind: when the federal regulator does it, 'everyone sources faces this way' becomes harder to rebut and easier to hide behind.
- Universities and agencies holding requestable face datasets face pressure to justify their collection and sharing terms, since the FT's catalogue shows the practice extends well beyond one benchmark program.
Third-order effects
- If the pattern holds, consent becomes the fault line that separates facial recognition governance debates from generic AI-ethics debates — pushing toward rules that treat biometric imagery as categorically different from other scraped data, with government collections held to the standard regulators demand of vendors.
- Benchmark programs themselves may need independent review of their data provenance, because whoever controls the test set controls which algorithms US agencies buy — a power that is only legitimate if the test data itself is defensible.
The trend: Biometric data is forcing a redrawn public-data permission boundary, where consent — not accuracy — becomes the deciding question for how face datasets are collected, shared, and used to rank commercial systems.