Analysis: Books3, a dataset used to train Meta's Llama, BloombergGPT, and EleutherAI's GPT-J, contains 170K+ books from authors like Stephen King and Junot Díaz
One of the most troubling issues around generative AI is simple: It's being made in secret.
The Atlantic Alex Reisner
Context & Ripple Effects
Books3’s identification as a training source makes the provenance of major language models a concrete issue rather than an abstract concern about opaque AI development. It connects named authors’ works to models from Meta, Bloomberg, and EleutherAI.
The coverage later tracks both attempts to remove Books3 from public circulation and research finding that leading models can reproduce long book excerpts under strategic prompting. Together, those developments sharpen the stakes around what training-data disclosure and control can accomplish after models have been trained.
First-order effects
- Authors whose books appear in Books3 gain a clearer factual basis to scrutinize how their work was used in training the named models.
- Meta, Bloomberg, and EleutherAI face more direct questions about dataset provenance and whether they can document or defend the use of Books3.
Second-order effects
- Training-data sourcing becomes a competitive and operational issue: developers using large, poorly documented corpora may face greater pressure to audit datasets or seek clearer permissions.
- Efforts to take down Books3 may reduce access for future model builders, while doing little to reverse use by organizations that had already obtained the corpus.
Third-order effects
- If disclosure continues to connect model behavior to particular copyrighted works, AI development is likely to shift toward governed corpora with clearer provenance, permissions, and audit trails.
- The episode highlights a potential divide between firms able to secure and document content rights and smaller developers reliant on broadly available datasets.
The trend: Generative AI is moving from opaque web-scale data collection toward a contest over provenance, permission, and control of the corpora behind model capabilities.
Related: Governed AI Corpus · Public-data permission boundary · AI content commercialization · Meta · Books3 training-data analysis · Efforts to remove Books3
Related Coverage
- Anti-Piracy Group Takes Prominent AI Training Dataset “Books3′ Offline TorrentFreak · Ernesto Van der Sar
- Quadrillions not billions — Some five months after I highlighted the role of pirated books … AI and Copyright · Peter Schoppert
- Anti-Piracy Group Takes Massive AI Training Dataset 'Books3′ Offline Gizmodo · Kyle Barr
- Books3, a huge literary dataset used to train OpenAI and Meta chatbots, is filled with copyrighted works. The Verge · Wes Davis
- A giant online book collection Meta used to train its AI is gone over copyright issues Mashable · Alex Perry
- Had a ton of debates around this clash of culture with friends and it's one of those things where there isn't a right answer. In this specific case, legal implications aside (which are very very interesting), it feels kind-of wrong that authors don't have any say as to how their works are used nor are they compensated for it. … @acookiecrumbles@indieweb.social
- Looking forward to when it transpires how all those publishers and analytics companies that sued sci-hub were also training their AI products on it. — https://www.theatlantic.com/ ... @Samuelmoore@hcommons.social · Samuel Moore
- On pirated books being used to train AI. I still think there's a good chance that a US court will find the use fair, or even transformative. I think EU courts may not. — https://www.theatlantic.com/ ... @mpe@ravenation.club · Martin Paul Eve
- The AI crookedness goes on. — Revealed: The Authors Whose Pirated Books Are Powering Generative AI — Stephen King, Zadie Smith, and Michael Pollan are among thousands of writers whose copyrighted works are being used to train large language models. — [the article is mostly behind a paywall but maybe someone knows a work around. … @johnshirley2024@wandering.shop · John Shirley
- We need robust and comprehensive systems for licensing and using datasets for ML training. It's not just about well known authors, it's about your data too. … João Fiadeiro
- The Authors Whose Pirated Books Are Powering Generative AI Hacker News
Discussion
-
@TryshHQ@mastodon.social
Jim Parsons
on mastodon
• #GenerativeAI remains a pipe dream — • the evil it's unleashed is 100% real — “Revealed: The Authors Whose Pirated Books Are Powering Generative AI” — https://www.theatlantic.com/ ... 1. Is there a better #SiliconValley #BigTech initiative to “flood the zone with shit” (…
-
@technursejon@mastodon.art
Jon
on mastodon
“The exploitation of pirated books for profit, with the goal of replacing the writers whose work was taken—this is a different and disturbing trend.” #AI #ArtificialIntelligence — https://www.theatlantic.com/ ...
-
@Dhmspector@mastodon.social
Dave Spector
on mastodon
I remember back in the #80s when the #FBI would kick in the front doors of the homes of #teens where were allegedly #prating #software. Ah! Good times. — I eagerly await(*) seeing various members of the #billionaire #brats #club like #SamAltman in cuffs charged with these mas…
-
@elkmovie@mastodon.social
Michael Love
on mastodon
It's starting to feel like - much as with crypto - generative AI is not going to become big enough fast enough to outrun the law. https://www.theatlantic.com/ ...
-
@nash076
@nash076
on x
The MO for every single one of these techbro operations is to just break any law that's in the way and dare someone to take them to court. And every time, the end result has just been a rickety facade that falls over at the same time they're running off with the money.
-
@cromwellian
@cromwellian
on x
I don't get why this is piracy or copyright infringement. If I read a book and gain knowledge or ideas from it and tell someone else in my own words, it's not theft. So why is an AI doing the same thing theft? Yes, if it overfits and regurgitates large sections of the book, sure.
-
@garymarcus
Gary Marcus
on x
@glynmoody These systems don't analyze the *meanings* of books; they inhale word sequences. It's different,
-
@glynmoody
Glyn Moody
on x
so when you read and analyse a book - which is what the AI systems did, not copy it - that's stealing? got it...
-
@cromwellian
@cromwellian
on x
I mean, didn't we already go through this with building indexes in search engines? Or people wanting to get paid for people linking to them. All of this friction does more to stall progress and in the end probably doesn't help the authors.
-
@dwcongdon
David W. Congdon
on x
Watch this get a pass while Internet Archive is shut down.
-
@galbeckerman
Gal Beckerman
on x
Just a sense of the scope here: “More than 30,000 titles are from Penguin Random House and its imprints, 14,000 from HarperCollins, 7,000 from Macmillan, 1,800 from Oxford University Press, and 600 from Verso.”
-
@lmatsakis
Louise Matsakis
on x
The Books3 dataset was hiding in the open, but no one bothered to analyze it. There are probably tons of others that similarly contain stolen works [image]
-
@avishaiw
@avishaiw
on x
“Pirated books are being used as inputs for computer programs that are changing how we read, learn, and communicate. The future promised by AI is written with stolen words.”
-
@dlberes
Damon Beres
on x
Books3 is, to an extent, a known quantity, especially after a recent takedown request. (You'll see plenty of references in the piece.) This story represents the first comprehensive analysis of its contents for people outside of the AI community.
-
@srnlrsn
@srnlrsn
on x
Update: AI is having a Napster moment. https://www.theatlantic.com/ ... [image]
-
@leifweatherby
@leifweatherby
on x
amazing piece Alex and @dlberes - super important observations late in the article about the balance between copyright and distribution. if Books3 is taken out of LlaMA that will be a boon to OpenAI, at least for the moment. data wars are fully underway https://www.theatlantic.co…
-
@iamrobotbear
@iamrobotbear
on x
@dlberes ... This is misleading as hell. Meta didn't release their LLaMA dataset first off. Secondly, look at the Google Books case and pls explain how this isn't transformative.
-
@garymarcus
Gary Marcus
on x
“The future promised by AI is written with stolen words.” Literally:
-
@marc__watkins
Marc Watkins
on x
To emphasize how FUBAR this whole situation is with training data that may have been pirated: [image]
-
@lmatsakis
Louise Matsakis
on x
It's really instructive how the dataset here was analyzed. To understand how generative AI was built, researchers and journalists are going to need to rely on similar techniques https://www.theatlantic.com/ ... [image]
-
@heidilegg
Heidi Legg
on x
First journalists, now authors. How Big Tech Platforms wiped out many essential workers by refusing to pay for our work.
-
@lmatsakis
Louise Matsakis
on x
Meta trained LLaMA on upwards of *170,000* pirated books, the majority of which were written in the last 20 years. Great scoop from @TheAtlantic https://www.theatlantic.com/ ...
-
@mkirschenbaum
Matthew Kirschenbaum
on x
This is a major, major piece of reporting and computer forensics at the intersect of #criticalAI, #bookhistory, and #dh. Hats off to @dlberes and of course the author, Alex Reisner. [image]
-
@stevesi
Steven Sinofsky
on x
Generative-AI models by Meta and others have been trained on a secret trove of more than 170,000 books to gain their power. In a new analysis, Alex Reisner uncovers the authors whose work has been fed into popular AI programs: https://www.theatlantic.com/ ...
-
@kbandersen
Kurt Andersen
on x
A corpus called Book3 used to train AIs by @Meta, @Bloomberg etc includes 170K+ post -2000 books, 1/3 fiction, 2/3 nonfiction. Thx for the piece, @TheAtlantic, but Alex Reisner is a programmer so please have them create an EZ search page for authors to see if our books were used.
-
@theatlantic
@theatlantic
on x
Generative-AI models by Meta and others have been trained on a secret trove of more than 170,000 books to gain their power. In a new analysis, Alex Reisner uncovers the authors whose work has been fed into popular AI programs: https://www.theatlantic.com/ ...
-
@sam_fentress
Sam Fentress
on x
you heard it here first: meta's LLM was trained on, among other things, 600 Verso titles https://www.theatlantic.com/ ...
-
@kz_howell
K. Z. Howell
on x
As expected, plagiarism writ large. Artificial intelligence is neither artificial nor intelligent, it is theft on a scale even government is incapable of. https://www.theatlantic.com/ ...
-
@dlberes
Damon Beres
on x
NEW: Meta, Bloomberg, and EleutherAI have trained generative AI on a dataset including upwards of 170,000 pirated books from authors like Stephen King, Zadie Smith, Margaret Atwood. Legality is complex. We have new details and context. tip @Techmeme https://www.theatlantic.com/ .…
-
@kylebarr5
Kyle Barr
on x
My latest for @Gizmodo hits on the cross-pollination of piracy ethics, copyright, and AI. Major companies like @MetaAI have trained their models on copyrighted works, but while tech giants can weather the storm of IP lawsuits and takedowns, small fry have a much harder time.
-
@lincodega
@lincodega
on x
Copyright is always a difficult part of the law to navigate, but adding AI training is what will give a lot of power back to authors and IP owners.@KyleBarr5 nails it here. [image]
-
@theshawwn
Shawn Presser
on x
A thoughtful Gizmodo article on books3 by @KyleBarr5: https://gizmodo.com/... Who would've guessed that the academictorrents website would be the last safe haven for AI research?
-
@jeffreygoldberg
Jeffrey Goldberg
on x
A story about book pirates and AI: https://www.theatlantic.com/ ...
-
@ivanthek
@ivanthek
on x
“The future promised by AI is written with stolen words.” The Achilles Heel of the current LLM craze. Lawyers will get paid. https://www.theatlantic.com/ ...
-
r/technology
r
on reddit
Revealed: The Authors Whose Pirated Books Are Powering Generative AI | Stephen King, Zadie Smith, and Michael Pollan are among thousands …