/
Navigation
Chronicles
Browse all articles
Explore
Semantic exploration
Research
Entity momentum
Nexus
Correlations & relationships
Story Arc
Topic evolution
Drift Map
Semantic trajectory animation
Posts
Analysis & commentary
Pulse API
Tech news intelligence API
Browse
Entities
Companies, people, products, technologies
Domains
Browse by publication source
Handles
Browse by social media handle
Detection
Concept Search
Semantic similarity search
High Impact Stories
Top coverage by position
Sentiment Analysis
Positive/negative coverage
Anomaly Detection
Unusual coverage patterns
Analysis
Rivalry Report
Compare two entities head-to-head
Semantic Pivots
Narrative discontinuities
Crisis Response
Event recovery patterns
Connected
Search: /
Command: ⌘K
Embeddings: large
TEXXR

Chronicles

The story behind the story

← → days · ↑ ↓ browse · Enter similar · o open

Inside LAION-5B, an AI training dataset of 5B+ images that has been unavailable for download after researchers found 3,000+ instances of CSAM in December 2023

If you want to make a really big AI model — the kind that can generate images or do your homework, or build this website …

Knowing Machines

Context & Ripple Effects

LAION-5B had become a widely used open training-data resource, including in work associated with Google’s Imagen and Stable Diffusion, after LAION’s free five-billion-image collection gained broad model-development use.

The download withdrawal follows researchers’ December finding of more than 1,000 CSAM instances, while this examination reports a larger tally. It turns a data-quality and safety failure into an access problem for downstream model builders.

First-order effects

  • LAION-5B is unavailable for download, removing a major open-source image-data input for researchers and developers that relied on the collection.
  • The discovery exposes the dataset’s curation controls to scrutiny and raises immediate safety and provenance concerns for organizations that used it in model training.

Second-order effects

  • Developers seeking comparable training data must place more weight on screening, documentation, and controlled access rather than treating a large public corpus as a ready-to-use input.
  • Open-data providers and model builders face pressure to make dataset safety checks auditable, especially where harmful material can be reproduced or amplified by generative systems.

Third-order effects

  • If large shared corpora increasingly require withdrawal or restricted distribution after safety failures, access to vetted training data may become a durable bottleneck that favors organizations able to fund data governance.
  • The episode points to AI commons being treated less as informal research infrastructure and more as critical infrastructure with ongoing accountability for its downstream use.

The trend: Generative AI’s expansion is making the provenance, safety review, and governance of shared training data as consequential as model architecture itself.

Discussion

  • @sethcjayson Seth Jayson on threads
    This is SUCH a great explanation of how the AI giants are building The Next Big Thing on a foundation of sand, with unknown quantities of rotten wood mixed in. https://knowingmachines.org/ ...
  • @ethanzuckerman Ethan Zuckerman on threads
    This is both terrific research and absolutely beautiful presentation of data.  Adding this to my “recommended readings” on AI and ethical implications - super smart and affecting work: https://knowingmachines.org/ ...
  • @renee.diresta Renee DiResta on threads
    Enjoyable and very informative visualized learning experience here related to #AI training data: https://knowingmachines.org/ ... “Here we find an important truth about LAION-5B: It contains less about how humans see the world than it does about how search engines see the world. …
  • @stungeye Kyle Geske on x
    The LAION dataset is almost entirely machine-curated using ML models, often in faulty or biased ways. Where human rankings are used, they come from a very small number of usually western folks with niche tastes.
  • @katecrawford Kate Crawford on x
    👉Today we're launching this investigation into LAION-5B, the blockbuster dataset behind Midjourney and Stable Diffusion. It's a deep dive into how the dataset was made, and where the images come from. The brilliant @christo_buschek & @blprnt follow the models all the way down.
  • @iethics @iethics on x
    Another “truth about generative #AI: The [concept] of what is... visually appealing can be influenced in outsized ways by the tastes of a very small group of individuals, and the processes that are chosen by dataset creators to curate the datasets” https://knowingmachines.org/ ..…
  • @brainpicker Maria Popova on x
    “Investigating training sets is an essential avenue to understanding how generative AI models work; the ways they see and re-create the world.” Models All the Way Down — fantastic (and haunting in its intimations) project by my friend @blprnt https://knowingmachines.org/ ...
  • @madeehamerchant Madeeha Merchant on x
    Ever wondered about the anatomy of GenAI ? Amazing work, X-raying the LAION-5B dataset ! Always a fan of @blprnt
  • @iethics @iethics on x
    “There are models on top of models, and trainings sets on top of training sets. Omissions and biases and blind spots from these stacked-up models and training sets shape all of the resulting new models and new training sets”: https://knowingmachines.org/ ... #ethics #data #AI #re…
  • @iethics @iethics on x
    “The tiniest of shifts in LAION's thresholds could have excluded or included hundreds of millions of images. What the images contain plays no role at all in deciding what stays and what goes.” https://knowingmachines.org/ ... #ethics #AI #data
  • @christo_buschek Christo Buschek on x
    The AI field's goal is nothing less than to transform the world. But what are the foundations upon which this transformation is built? In this investigation, @blprnt and I looked at LAION-5B, the only open-source foundation dataset currently available. [image]
  • @mattocko Matt Ocko on x
    Visually impactful & thoughtful piece on the serious problem of recursive bad data poisoning of large models' training data sets — and hence the models themselves https://knowingmachines.org/ ... H/t @mattgreenfield