Inside LAION-5B, an AI training dataset of 5B+ images that has been unavailable for download after researchers found 3,000+ instances of CSAM in December 2023
If you want to make a really big AI model — the kind that can generate images or do your homework, or build this website …
LAION-5B is unavailable for download, removing a major open-source image-data input for researchers and developers that relied on the collection.
The discovery exposes the dataset’s curation controls to scrutiny and raises immediate safety and provenance concerns for organizations that used it in model training.
Second-order effects
Developers seeking comparable training data must place more weight on screening, documentation, and controlled access rather than treating a large public corpus as a ready-to-use input.
Open-data providers and model builders face pressure to make dataset safety checks auditable, especially where harmful material can be reproduced or amplified by generative systems.
Third-order effects
If large shared corpora increasingly require withdrawal or restricted distribution after safety failures, access to vetted training data may become a durable bottleneck that favors organizations able to fund data governance.
The episode points to AI commons being treated less as informal research infrastructure and more as critical infrastructure with ongoing accountability for its downstream use.
The trend:Generative AI’s expansion is making the provenance, safety review, and governance of shared training data as consequential as model architecture itself.
This is SUCH a great explanation of how the AI giants are building The Next Big Thing on a foundation of sand, with unknown quantities of rotten wood mixed in. https://knowingmachines.org/ ...
This is both terrific research and absolutely beautiful presentation of data. Adding this to my “recommended readings” on AI and ethical implications - super smart and affecting work: https://knowingmachines.org/ ...
Enjoyable and very informative visualized learning experience here related to #AI training data: https://knowingmachines.org/ ... “Here we find an important truth about LAION-5B: It contains less about how humans see the world than it does about how search engines see the world. …
The LAION dataset is almost entirely machine-curated using ML models, often in faulty or biased ways. Where human rankings are used, they come from a very small number of usually western folks with niche tastes.
👉Today we're launching this investigation into LAION-5B, the blockbuster dataset behind Midjourney and Stable Diffusion. It's a deep dive into how the dataset was made, and where the images come from. The brilliant @christo_buschek & @blprnt follow the models all the way down.
Another “truth about generative #AI: The [concept] of what is... visually appealing can be influenced in outsized ways by the tastes of a very small group of individuals, and the processes that are chosen by dataset creators to curate the datasets” https://knowingmachines.org/ ..…
“Investigating training sets is an essential avenue to understanding how generative AI models work; the ways they see and re-create the world.” Models All the Way Down — fantastic (and haunting in its intimations) project by my friend @blprnt https://knowingmachines.org/ ...
“There are models on top of models, and trainings sets on top of training sets. Omissions and biases and blind spots from these stacked-up models and training sets shape all of the resulting new models and new training sets”: https://knowingmachines.org/ ... #ethics #data #AI #re…
“The tiniest of shifts in LAION's thresholds could have excluded or included hundreds of millions of images. What the images contain plays no role at all in deciding what stays and what goes.” https://knowingmachines.org/ ... #ethics #AI #data
The AI field's goal is nothing less than to transform the world. But what are the foundations upon which this transformation is built? In this investigation, @blprnt and I looked at LAION-5B, the only open-source foundation dataset currently available. [image]
Visually impactful & thoughtful piece on the serious problem of recursive bad data poisoning of large models' training data sets — and hence the models themselves https://knowingmachines.org/ ... H/t @mattgreenfield