Experts say AI tools lack the training data needed to understand the cultural meaning of images, recipes, and other data concerning underrepresented cultures
W. Victor H. Yarlott, A.I. Researcher at FIU and member of the Crow Tribe of Montana https://www.nytimes.com/... Eli Chen / @storiesbyeli : I often tell people that I have the coolest boss and this is just one of many things that make @DavarArdalan such a stellar person to work for “You're not really representing human intelligence or human knowledge unless your system can handle it from a broad range of cultures.” https://twitter.com/... Nikki McLay / @aigrrl : The NY Times has featured our work! Almost 4 years ago we @ivowai came together to try and innovate around one of the most challenging aspects of A.I. - authentic inclusion of culture. As we continue into 2022, we continue to make progress... https://www.nytimes.com/... @nytimes : Artificial intelligence tools struggle to identify information involving Indigenous cultures and Native communities, often producing inaccurate results. Now some are pushing for underrepresented groups to create their own data. https://www.nytimes.com/... Dan Saltzstein / @dansaltzstein : I found this story fascinating: datasets (feeding AI-based image identification, etc.) on Indigenous culture are meager. A group is trying to change that. Can they? https://www.nytimes.com/...
Context & Ripple Effects
This 2022 Times piece gave a name to a gap that earlier coverage had only measured indirectly: the 21%-35% facial-recognition error rate for dark-skinned women at Microsoft, IBM, and Megvii was the symptom, and the missing cultural training data is the diagnosis. W. Victor H. Yarlott of FIU and the Crow Tribe, along with Nikki McLay's ivowai, argue that systems trained without underrepresented cultures' images, recipes, and knowledge simply cannot represent human intelligence broadly.
The arc since has run toward communities building the data themselves rather than waiting to be included: Indian nonprofit Karya sells training data while its workers keep ownership and profit, and Indigenous engineers are now building speech recognition for hundreds of endangered Native languages.
First-order effects
- AI tools misidentify Indigenous cultures and Native communities and return inaccurate results today, which is the direct cost Yarlott and other experts attribute to absent training data.
- Advocates like ivowai gain a concrete agenda: push underrepresented groups to create their own training data instead of relying on web-scraped corpora that omit them.
Second-order effects
- Data creation becomes an economic question, not just a technical one — Karya's model shows the alternative structure, where workers retain ownership of the data they make and capture the profit that platforms would otherwise take.
- As conventional web datasets run dry and flawed ones face scrutiny — LAION-5B was pulled from download after researchers found thousands of CSAM instances — curated, consented, culturally specific datasets rise in value precisely where general scraping fails.
Third-order effects
- If companies are forced toward smaller, more specialized models as general-purpose data exhausts, community-owned datasets become strategic inputs, shifting bargaining power toward the groups whose knowledge was previously taken for free.
- Dataset governance moves toward provenance and consent standards, with the LAION-5B takedown and Karya's ownership model marking the two poles: unvetted scraping versus compensated, rights-retaining contribution.
The trend: Training data is migrating from free-for-all web scraping toward community-owned, purpose-built datasets, driven by both cultural accuracy failures and the exhaustion of conventional sources.