Inside LAION-5B, an AI training dataset of 5B+ images that has been unavailable for download after researchers found 3,000+ instances of CSAM in December 2023
All — The — Way — Down — This one is sooo good. I recommend this to anyone playing with #AI to understand the biases and the complexities. Oh and the discussion of alt text is amazing. … Erik Jonker / @ErikJonker@mastodon.social : Just wow...amazing website/visualization about LAION-5B , a large dataset a lot models are trained on. — https://knowingmachines.org/ ... #AI #bigdata #LAION5B #trainingdata #CSAM Larry O'Brien / @Lobrien@hachyderm.io : Absolutely wonderful article about how datasets are profoundly influenced by near-arbitrary magic numbers, self-selecting human reviewers, and WEIRD over-representation. https://knowingmachines.org/ ... #ML Bastian Greshake Tzovaras / @gedankenstuecke … : «Models All The Way Down» is a really interesting exploration of how an image data set (in this case LAION-5B) was assembled to be used for training ML/"AI" models! — > “It contains less about how humans see the world than it does about how search engines see the world. … X: Ethan Zuckerman / @ethanz : This is both terrific research and absolutely beautiful presentation of data. Adding this to my “recommended readings” on AI and ethical implications - super smart and affecting work. Kyle Geske / @stungeye : The LAION dataset is almost entirely machine-curated using ML models, often in faulty or biased ways. Where human rankings are used, they come from a very small number of usually western folks with niche tastes. Maria Popova / @brainpicker : “Investigating training sets is an essential avenue to understanding how generative AI models work; the ways they see and re-create the world.” Models All the Way Down — fantastic (and haunting in its intimations) project by my friend @blprnt https://knowingmachines.org/ ... @iethics : “There are models on top of models, and trainings sets on top of training sets. Omissions and biases and blind spots from these stacked-up models and training sets shape all of the resulting new models and new training sets”: https://knowingmachines.org/ ... #ethics #data #AI #research @iethics : Another “truth about generative #AI: The [concept] of what is... visually appealing can be influenced in outsized ways by the tastes of a very small group of individuals, and the processes that are chosen by dataset creators to curate the datasets” https://knowingmachines.org/ ... #data @iethics : “The tiniest of shifts in LAION's thresholds could have excluded or included hundreds of millions of images. What the images contain plays no role at all in deciding what stays and what goes.” https://knowingmachines.org/ ... #ethics #AI #data Matt Ocko / @mattocko : Visually impactful & thoughtful piece on the serious problem of recursive bad data poisoning of large models' training data sets — and hence the models themselves https://knowingmachines.org/ ... H/t @mattgreenfield Kate Crawford / @katecrawford : 👉Today we're launching this investigation into LAION-5B, the blockbuster dataset behind Midjourney and Stable Diffusion. It's a deep dive into how the dataset was made, and where the images come from. The brilliant @christo_buschek & @blprnt follow the models all the way down. Madeeha Merchant / @madeehamerchant : Ever wondered about the anatomy of GenAI ? Amazing work, X-raying the LAION-5B dataset ! Always a fan of @blprnt Christo Buschek / @christo_buschek : The AI field's goal is nothing less than to transform the world. But what are the foundations upon which this transformation is built? In this investigation, @blprnt and I looked at LAION-5B, the only open-source foundation dataset currently available. [image] LinkedIn: Matt Greenfield : This is one of the best explanations I have seen of the way that small groups of humans and arbitrary selection criteria create not just bias … Forums: Hacker News : Models all the way down
Context & Ripple Effects
LAION-5B had already become a consequential shared input for image-generation systems, while earlier analysis found that a large share of Stable Diffusion’s training images came from a narrow set of domains, underscoring how upstream collection choices concentrate downstream influence. the concentration of source domains in Stable Diffusion’s training data made the provenance of a supposedly broad web-scale corpus material.
The dataset’s removal follows researchers’ earlier identification of CSAM in LAION-5B. This investigation adds detail on how machine-led curation, limited human ranking, and threshold choices can embed omissions and bias before models are trained. Earlier researchers had documented CSAM in the corpus
First-order effects
- LAION-5B remains unavailable for download, removing a major open dataset option for groups seeking to train large image models from a common corpus.
- Developers of systems trained on or associated with LAION-5B, including Midjourney and Stable Diffusion, face sharper scrutiny of training-data provenance, safety exposure, and inherited bias.
Second-order effects
- Open-model developers and dataset maintainers are pushed toward more auditable filtering, documentation, and human review; the account shows that small selection-threshold changes can alter the corpus at enormous scale.
- Demand rises for specialized data-review work and safety controls, extending the reliance on the often-hidden annotation labor described in the AI data-labeling workforce rather than treating automated curation as sufficient.
Third-order effects
- If large shared datasets cannot be safely maintained as open downloads, access to training-ready data may consolidate among organizations able to fund collection, review, and compliance—turning the AI commons into critical infrastructure with higher operating barriers.
- The episode strengthens the case that model behavior must be evaluated as a product of layered data pipelines, not just model architecture; whether this yields common audit standards depends on how dataset builders and model developers respond.
The trend: Generative AI is moving from treating web-scale training data as an abundant commodity toward treating its provenance, safety, and governance as core model infrastructure.