Analysis: Books3, a dataset used to train Meta's Llama, BloombergGPT, and EleutherAI's GPT-J, contains 170K+ books from Stephen King and other authors
One of the most troubling issues around generative AI is simple: It's being made in secret. To produce humanlike answers to questions …
The AtlanticAlex Reisner
Context & Ripple Effects
The Books3 disclosure put named commercial and open-model projects in the same training-data controversy, making the provenance of literary corpora a concrete issue rather than an abstract concern. Follow-on coverage showed copyright activists seeking to remove the dataset from circulation through a campaign to wipe Books3 from the internet.
The issue fits a broader pipeline: later reporting examined LibGen as a major pirate-library source allegedly used for AI training, while researchers later found that leading models could produce long book excerpts under strategic prompting. Together, those developments sharpen the distinction between access to text for training and control over how models may reproduce it.
First-order effects
Authors and rightsholders whose books appear in Books3 gain a clearer factual basis to scrutinize the training practices behind Meta's Llama, BloombergGPT, and GPT-J.
The named model developers face immediate provenance and reputational pressure because a dataset containing more than 170,000 books is publicly tied to their systems.
Second-order effects
Efforts to remove Books3 may reduce easy access for later entrants, while companies that already trained on it retain the resulting model capabilities—an asymmetry highlighted by the push to take Books3 offline.
Publishers and AI developers have stronger incentives to distinguish licensed or documented text sources from corpora associated with pirate-library distribution, including the LibGen training-data allegations.
Third-order effects
If disclosures and reproducibility findings continue to accumulate, training-data provenance is likely to become a competitive and governance constraint alongside model quality, especially for systems built on cultural works.
The industry may move toward more explicit licensing, auditing, and output safeguards; whether that narrows the gap between established model owners and smaller developers depends on how access rules are applied.
The trend:Generative AI is shifting from opaque web-scale data collection toward a contested market for traceable, governable rights to high-value training content.
“The exploitation of pirated books for profit, with the goal of replacing the writers whose work was taken—this is a different and disturbing trend.” #AI #ArtificialIntelligence — https://www.theatlantic.com/ ...
This is a major, major piece of reporting and computer forensics at the intersect of #criticalAI, #bookhistory, and #dh. Hats off to @dlberes and of course the author, Alex Reisner. [image]
A corpus called Book3 used to train AIs by @Meta, @Bloomberg etc includes 170K+ post -2000 books, 1/3 fiction, 2/3 nonfiction. Thx for the piece, @TheAtlantic, but Alex Reisner is a programmer so please have them create an EZ search page for authors to see if our books were used.
Generative-AI models by Meta and others have been trained on a secret trove of more than 170,000 books to gain their power. In a new analysis, Alex Reisner uncovers the authors whose work has been fed into popular AI programs: https://www.theatlantic.com/ ...
NEW: Meta, Bloomberg, and EleutherAI have trained generative AI on a dataset including upwards of 170,000 pirated books from authors like Stephen King, Zadie Smith, Margaret Atwood. Legality is complex. We have new details and context. tip @Techmeme https://www.theatlantic.com/ .…
A thoughtful Gizmodo article on books3 by @KyleBarr5: https://gizmodo.com/... Who would've guessed that the academictorrents website would be the last safe haven for AI research?
Generative-AI models by Meta and others have been trained on a secret trove of more than 170,000 books to gain their power. In a new analysis, Alex Reisner uncovers the authors whose work has been fed into popular AI programs: https://www.theatlantic.com/ ...
My latest for @Gizmodo hits on the cross-pollination of piracy ethics, copyright, and AI. Major companies like @MetaAI have trained their models on copyrighted works, but while tech giants can weather the storm of IP lawsuits and takedowns, small fry have a much harder time.
Copyright is always a difficult part of the law to navigate, but adding AI training is what will give a lot of power back to authors and IP owners.@KyleBarr5 nails it here. [image]
“The future promised by AI is written with stolen words.” The Achilles Heel of the current LLM craze. Lawyers will get paid. https://www.theatlantic.com/ ...
As expected, plagiarism writ large. Artificial intelligence is neither artificial nor intelligent, it is theft on a scale even government is incapable of. https://www.theatlantic.com/ ...