/
Navigation
Chronicles
Browse all articles
Explore
Semantic exploration
Research
Entity momentum
Nexus
Correlations & relationships
Story Arc
Topic evolution
Drift Map
Semantic trajectory animation
Posts
Analysis & commentary
Pulse API
Tech news intelligence API
Browse
Entities
Companies, people, products, technologies
Domains
Browse by publication source
Handles
Browse by social media handle
Detection
Concept Search
Semantic similarity search
High Impact Stories
Top coverage by position
Sentiment Analysis
Positive/negative coverage
Anomaly Detection
Unusual coverage patterns
Analysis
Rivalry Report
Compare two entities head-to-head
Semantic Pivots
Narrative discontinuities
Crisis Response
Event recovery patterns
Connected
Search: /
Command: ⌘K
Embeddings: large
TEXXR

Chronicles

The story behind the story

days · browse · Enter similar · o open

Experts say AI tools lack the training data needed to understand the cultural meaning of images, recipes, and other data concerning underrepresented cultures

W. Victor H. Yarlott, A.I. Researcher at FIU and member of the Crow Tribe of Montana https://www.nytimes.com/... Eli Chen / @storiesbyeli : I often tell people that I have the coolest boss and this is just one of many things that make @DavarArdalan such a stellar person to work for “You're not really representing human intelligence or human knowledge unless your system can handle it from a broad range of cultures.” https://twitter.com/... Nikki McLay / @aigrrl : The NY Times has featured our work! Almost 4 years ago we @ivowai came together to try and innovate around one of the most challenging aspects of A.I. - authentic inclusion of culture. As we continue into 2022, we continue to make progress... https://www.nytimes.com/... @nytimes : Artificial intelligence tools struggle to identify information involving Indigenous cultures and Native communities, often producing inaccurate results. Now some are pushing for underrepresented groups to create their own data. https://www.nytimes.com/... Dan Saltzstein / @dansaltzstein : I found this story fascinating: datasets (feeding AI-based image identification, etc.) on Indigenous culture are meager. A group is trying to change that. Can they? https://www.nytimes.com/...

New York Times Alex V. Cipolle

Context & Ripple Effects

This 2022 Times piece gave a name to a gap that earlier coverage had only measured indirectly: the 21%-35% facial-recognition error rate for dark-skinned women at Microsoft, IBM, and Megvii was the symptom, and the missing cultural training data is the diagnosis. W. Victor H. Yarlott of FIU and the Crow Tribe, along with Nikki McLay's ivowai, argue that systems trained without underrepresented cultures' images, recipes, and knowledge simply cannot represent human intelligence broadly.

The arc since has run toward communities building the data themselves rather than waiting to be included: Indian nonprofit Karya sells training data while its workers keep ownership and profit, and Indigenous engineers are now building speech recognition for hundreds of endangered Native languages.

First-order effects

  • AI tools misidentify Indigenous cultures and Native communities and return inaccurate results today, which is the direct cost Yarlott and other experts attribute to absent training data.
  • Advocates like ivowai gain a concrete agenda: push underrepresented groups to create their own training data instead of relying on web-scraped corpora that omit them.

Second-order effects

  • Data creation becomes an economic question, not just a technical one — Karya's model shows the alternative structure, where workers retain ownership of the data they make and capture the profit that platforms would otherwise take.
  • As conventional web datasets run dry and flawed ones face scrutiny — LAION-5B was pulled from download after researchers found thousands of CSAM instances — curated, consented, culturally specific datasets rise in value precisely where general scraping fails.

Third-order effects

  • If companies are forced toward smaller, more specialized models as general-purpose data exhausts, community-owned datasets become strategic inputs, shifting bargaining power toward the groups whose knowledge was previously taken for free.
  • Dataset governance moves toward provenance and consent standards, with the LAION-5B takedown and Karya's ownership model marking the two poles: unvetted scraping versus compensated, rights-retaining contribution.

The trend: Training data is migrating from free-for-all web scraping toward community-owned, purpose-built datasets, driven by both cultural accuracy failures and the exhaustion of conventional sources.

Discussion

  • @nytimestech @nytimestech on x
    Image identification and other A.I. tools are terrible at tagging information involving Indigenous culture. Is there a solution? https://www.nytimes.com/...
  • @katrina_hrm Katrina Jones on x
    “You're not really representing human intelligence or human knowledge unless your system can handle it from a broad range of cultures.” — W. Victor H. Yarlott, A.I. Researcher at FIU and member of the Crow Tribe of Montana https://www.nytimes.com/...
  • @storiesbyeli Eli Chen on x
    I often tell people that I have the coolest boss and this is just one of many things that make @DavarArdalan such a stellar person to work for “You're not really representing human intelligence or human knowledge unless your system can handle it from a broad range of cultures.” h…
  • @aigrrl Nikki McLay on x
    The NY Times has featured our work! Almost 4 years ago we @ivowai came together to try and innovate around one of the most challenging aspects of A.I. - authentic inclusion of culture. As we continue into 2022, we continue to make progress... https://www.nytimes.com/...
  • @nytimes @nytimes on x
    Artificial intelligence tools struggle to identify information involving Indigenous cultures and Native communities, often producing inaccurate results. Now some are pushing for underrepresented groups to create their own data. https://www.nytimes.com/...
  • @dansaltzstein Dan Saltzstein on x
    I found this story fascinating: datasets (feeding AI-based image identification, etc.) on Indigenous culture are meager. A group is trying to change that. Can they? https://www.nytimes.com/...