/
Navigation
Chronicles
Browse all articles
Explore
Semantic exploration
Research
Entity momentum
Nexus
Correlations & relationships
Story Arc
Topic evolution
Drift Map
Semantic trajectory animation
Posts
Analysis & commentary
Pulse API
Tech news intelligence API
Browse
Entities
Companies, people, products, technologies
Domains
Browse by publication source
Handles
Browse by social media handle
Detection
Concept Search
Semantic similarity search
High Impact Stories
Top coverage by position
Sentiment Analysis
Positive/negative coverage
Anomaly Detection
Unusual coverage patterns
Analysis
Rivalry Report
Compare two entities head-to-head
Semantic Pivots
Narrative discontinuities
Crisis Response
Event recovery patterns
Connected
Search: /
Command: ⌘K
Embeddings: large
TEXXR

Chronicles

The story behind the story

days · browse · Enter similar · o open

Inside LAION-5B, an AI training dataset of 5B+ images that has been unavailable for download after researchers found 3,000+ instances of CSAM in December 2023

All  —  The  —  Way  —  Down  —  This one is sooo good.  I recommend this to anyone playing with #AI to understand the biases and the complexities.  Oh and the discussion of alt text is amazing. … Erik Jonker / @ErikJonker@mastodon.social : Just wow...amazing website/visualization about LAION-5B , a large dataset a lot models are trained on.  —  https://knowingmachines.org/ ...  #AI #bigdata #LAION5B #trainingdata #CSAM Larry O'Brien / @Lobrien@hachyderm.io : Absolutely wonderful article about how datasets are profoundly influenced by near-arbitrary magic numbers, self-selecting human reviewers, and WEIRD over-representation. https://knowingmachines.org/ ... #ML Bastian Greshake Tzovaras / @gedankenstuecke … : «Models All The Way Down» is a really interesting exploration of how an image data set (in this case LAION-5B) was assembled to be used for training ML/"AI" models!  —  > “It contains less about how humans see the world than it does about how search engines see the world. … X: Ethan Zuckerman / @ethanz : This is both terrific research and absolutely beautiful presentation of data. Adding this to my “recommended readings” on AI and ethical implications - super smart and affecting work. Kyle Geske / @stungeye : The LAION dataset is almost entirely machine-curated using ML models, often in faulty or biased ways. Where human rankings are used, they come from a very small number of usually western folks with niche tastes. Maria Popova / @brainpicker : “Investigating training sets is an essential avenue to understanding how generative AI models work; the ways they see and re-create the world.” Models All the Way Down — fantastic (and haunting in its intimations) project by my friend @blprnt https://knowingmachines.org/ ... @iethics : “There are models on top of models, and trainings sets on top of training sets. Omissions and biases and blind spots from these stacked-up models and training sets shape all of the resulting new models and new training sets”: https://knowingmachines.org/ ... #ethics #data #AI #research @iethics : Another “truth about generative #AI: The [concept] of what is... visually appealing can be influenced in outsized ways by the tastes of a very small group of individuals, and the processes that are chosen by dataset creators to curate the datasets” https://knowingmachines.org/ ... #data @iethics : “The tiniest of shifts in LAION's thresholds could have excluded or included hundreds of millions of images. What the images contain plays no role at all in deciding what stays and what goes.” https://knowingmachines.org/ ... #ethics #AI #data Matt Ocko / @mattocko : Visually impactful & thoughtful piece on the serious problem of recursive bad data poisoning of large models' training data sets — and hence the models themselves https://knowingmachines.org/ ... H/t @mattgreenfield Kate Crawford / @katecrawford : 👉Today we're launching this investigation into LAION-5B, the blockbuster dataset behind Midjourney and Stable Diffusion. It's a deep dive into how the dataset was made, and where the images come from. The brilliant @christo_buschek & @blprnt follow the models all the way down. Madeeha Merchant / @madeehamerchant : Ever wondered about the anatomy of GenAI ? Amazing work, X-raying the LAION-5B dataset ! Always a fan of @blprnt Christo Buschek / @christo_buschek : The AI field's goal is nothing less than to transform the world. But what are the foundations upon which this transformation is built? In this investigation, @blprnt and I looked at LAION-5B, the only open-source foundation dataset currently available. [image] LinkedIn: Matt Greenfield : This is one of the best explanations I have seen of the way that small groups of humans and arbitrary selection criteria create not just bias … Forums: Hacker News : Models all the way down

Knowing Machines

Context & Ripple Effects

LAION-5B had already become a consequential shared input for image-generation systems, while earlier analysis found that a large share of Stable Diffusion’s training images came from a narrow set of domains, underscoring how upstream collection choices concentrate downstream influence. the concentration of source domains in Stable Diffusion’s training data made the provenance of a supposedly broad web-scale corpus material.

The dataset’s removal follows researchers’ earlier identification of CSAM in LAION-5B. This investigation adds detail on how machine-led curation, limited human ranking, and threshold choices can embed omissions and bias before models are trained. Earlier researchers had documented CSAM in the corpus

First-order effects

  • LAION-5B remains unavailable for download, removing a major open dataset option for groups seeking to train large image models from a common corpus.
  • Developers of systems trained on or associated with LAION-5B, including Midjourney and Stable Diffusion, face sharper scrutiny of training-data provenance, safety exposure, and inherited bias.

Second-order effects

  • Open-model developers and dataset maintainers are pushed toward more auditable filtering, documentation, and human review; the account shows that small selection-threshold changes can alter the corpus at enormous scale.
  • Demand rises for specialized data-review work and safety controls, extending the reliance on the often-hidden annotation labor described in the AI data-labeling workforce rather than treating automated curation as sufficient.

Third-order effects

  • If large shared datasets cannot be safely maintained as open downloads, access to training-ready data may consolidate among organizations able to fund collection, review, and compliance—turning the AI commons into critical infrastructure with higher operating barriers.
  • The episode strengthens the case that model behavior must be evaluated as a product of layered data pipelines, not just model architecture; whether this yields common audit standards depends on how dataset builders and model developers respond.

The trend: Generative AI is moving from treating web-scale training data as an abundant commodity toward treating its provenance, safety, and governance as core model infrastructure.

Discussion

  • @sethcjayson Seth Jayson on threads
    This is SUCH a great explanation of how the AI giants are building The Next Big Thing on a foundation of sand, with unknown quantities of rotten wood mixed in. https://knowingmachines.org/ ...
  • @renee.diresta Renee DiResta on threads
    Enjoyable and very informative visualized learning experience here related to #AI training data: https://knowingmachines.org/ ... “Here we find an important truth about LAION-5B: It contains less about how humans see the world than it does about how search engines see the world. …
  • @ethanzuckerman Ethan Zuckerman on threads
    This is both terrific research and absolutely beautiful presentation of data.  Adding this to my “recommended readings” on AI and ethical implications - super smart and affecting work: https://knowingmachines.org/ ...
  • @ErikJonker@mastodon.social Erik Jonker on mastodon
    Just wow...amazing website/visualization about LAION-5B , a large dataset a lot models are trained on.  —  https://knowingmachines.org/ ...  #AI #bigdata #LAION5B #trainingdata #CSAM
  • @ethanz Ethan Zuckerman on x
    This is both terrific research and absolutely beautiful presentation of data. Adding this to my “recommended readings” on AI and ethical implications - super smart and affecting work.
  • @stungeye Kyle Geske on x
    The LAION dataset is almost entirely machine-curated using ML models, often in faulty or biased ways. Where human rankings are used, they come from a very small number of usually western folks with niche tastes.
  • @brainpicker Maria Popova on x
    “Investigating training sets is an essential avenue to understanding how generative AI models work; the ways they see and re-create the world.” Models All the Way Down — fantastic (and haunting in its intimations) project by my friend @blprnt https://knowingmachines.org/ ...
  • @iethics @iethics on x
    “There are models on top of models, and trainings sets on top of training sets. Omissions and biases and blind spots from these stacked-up models and training sets shape all of the resulting new models and new training sets”: https://knowingmachines.org/ ... #ethics #data #AI #re…
  • @iethics @iethics on x
    Another “truth about generative #AI: The [concept] of what is... visually appealing can be influenced in outsized ways by the tastes of a very small group of individuals, and the processes that are chosen by dataset creators to curate the datasets” https://knowingmachines.org/ ..…
  • @iethics @iethics on x
    “The tiniest of shifts in LAION's thresholds could have excluded or included hundreds of millions of images. What the images contain plays no role at all in deciding what stays and what goes.” https://knowingmachines.org/ ... #ethics #AI #data
  • @mattocko Matt Ocko on x
    Visually impactful & thoughtful piece on the serious problem of recursive bad data poisoning of large models' training data sets — and hence the models themselves https://knowingmachines.org/ ... H/t @mattgreenfield
  • @katecrawford Kate Crawford on x
    👉Today we're launching this investigation into LAION-5B, the blockbuster dataset behind Midjourney and Stable Diffusion. It's a deep dive into how the dataset was made, and where the images come from. The brilliant @christo_buschek & @blprnt follow the models all the way down.
  • @madeehamerchant Madeeha Merchant on x
    Ever wondered about the anatomy of GenAI ? Amazing work, X-raying the LAION-5B dataset ! Always a fan of @blprnt
  • @christo_buschek Christo Buschek on x
    The AI field's goal is nothing less than to transform the world. But what are the foundations upon which this transformation is built? In this investigation, @blprnt and I looked at LAION-5B, the only open-source foundation dataset currently available. [image]