/
Navigation
Chronicles
Browse all articles
Explore
Semantic exploration
Research
Entity momentum
Nexus
Correlations & relationships
Story Arc
Topic evolution
Drift Map
Semantic trajectory animation
Posts
Analysis & commentary
Pulse API
Tech news intelligence API
Browse
Entities
Companies, people, products, technologies
Domains
Browse by publication source
Handles
Browse by social media handle
Detection
Concept Search
Semantic similarity search
High Impact Stories
Top coverage by position
Sentiment Analysis
Positive/negative coverage
Anomaly Detection
Unusual coverage patterns
Analysis
Rivalry Report
Compare two entities head-to-head
Semantic Pivots
Narrative discontinuities
Crisis Response
Event recovery patterns
Connected
Search: /
Command: ⌘K
Embeddings: large
TEXXR

Chronicles

The story behind the story

days · browse · Enter similar · o open

Analysis: Books3, a dataset used to train Meta's Llama, BloombergGPT, and EleutherAI's GPT-J, contains 170K+ books from authors like Stephen King and Junot Díaz

One of the most troubling issues around generative AI is simple: It's being made in secret.

The Atlantic Alex Reisner

Context & Ripple Effects

Books3’s identification as a training source makes the provenance of major language models a concrete issue rather than an abstract concern about opaque AI development. It connects named authors’ works to models from Meta, Bloomberg, and EleutherAI.

The coverage later tracks both attempts to remove Books3 from public circulation and research finding that leading models can reproduce long book excerpts under strategic prompting. Together, those developments sharpen the stakes around what training-data disclosure and control can accomplish after models have been trained.

First-order effects

  • Authors whose books appear in Books3 gain a clearer factual basis to scrutinize how their work was used in training the named models.
  • Meta, Bloomberg, and EleutherAI face more direct questions about dataset provenance and whether they can document or defend the use of Books3.

Second-order effects

  • Training-data sourcing becomes a competitive and operational issue: developers using large, poorly documented corpora may face greater pressure to audit datasets or seek clearer permissions.
  • Efforts to take down Books3 may reduce access for future model builders, while doing little to reverse use by organizations that had already obtained the corpus.

Third-order effects

  • If disclosure continues to connect model behavior to particular copyrighted works, AI development is likely to shift toward governed corpora with clearer provenance, permissions, and audit trails.
  • The episode highlights a potential divide between firms able to secure and document content rights and smaller developers reliant on broadly available datasets.

The trend: Generative AI is moving from opaque web-scale data collection toward a contest over provenance, permission, and control of the corpora behind model capabilities.

Discussion

  • @TryshHQ@mastodon.social Jim Parsons on mastodon
    • #GenerativeAI remains a pipe dream  —  • the evil it's unleashed is 100% real  —  “Revealed: The Authors Whose Pirated Books Are Powering Generative AI”  —  https://www.theatlantic.com/ ...  1. Is there a better #SiliconValley #BigTech initiative to “flood the zone with shit” (…
  • @technursejon@mastodon.art Jon on mastodon
    “The exploitation of pirated books for profit, with the goal of replacing the writers whose work was taken—this is a different and disturbing trend.”  #AI #ArtificialIntelligence  —  https://www.theatlantic.com/ ...
  • @Dhmspector@mastodon.social Dave Spector on mastodon
    I remember back in the #80s when the #FBI would kick in the front doors of the homes of #teens where were allegedly #prating #software.  Ah!  Good times.  —  I eagerly await(*) seeing various members of the #billionaire #brats #club like #SamAltman in cuffs charged with these mas…
  • @elkmovie@mastodon.social Michael Love on mastodon
    It's starting to feel like - much as with crypto - generative AI is not going to become big enough fast enough to outrun the law. https://www.theatlantic.com/ ...
  • @nash076 @nash076 on x
    The MO for every single one of these techbro operations is to just break any law that's in the way and dare someone to take them to court. And every time, the end result has just been a rickety facade that falls over at the same time they're running off with the money.
  • @cromwellian @cromwellian on x
    I don't get why this is piracy or copyright infringement. If I read a book and gain knowledge or ideas from it and tell someone else in my own words, it's not theft. So why is an AI doing the same thing theft? Yes, if it overfits and regurgitates large sections of the book, sure.
  • @garymarcus Gary Marcus on x
    @glynmoody These systems don't analyze the *meanings* of books; they inhale word sequences. It's different,
  • @glynmoody Glyn Moody on x
    so when you read and analyse a book - which is what the AI systems did, not copy it - that's stealing? got it...
  • @cromwellian @cromwellian on x
    I mean, didn't we already go through this with building indexes in search engines? Or people wanting to get paid for people linking to them. All of this friction does more to stall progress and in the end probably doesn't help the authors.
  • @dwcongdon David W. Congdon on x
    Watch this get a pass while Internet Archive is shut down.
  • @galbeckerman Gal Beckerman on x
    Just a sense of the scope here: “More than 30,000 titles are from Penguin Random House and its imprints, 14,000 from HarperCollins, 7,000 from Macmillan, 1,800 from Oxford University Press, and 600 from Verso.”
  • @lmatsakis Louise Matsakis on x
    The Books3 dataset was hiding in the open, but no one bothered to analyze it. There are probably tons of others that similarly contain stolen works [image]
  • @avishaiw @avishaiw on x
    “Pirated books are being used as inputs for computer programs that are changing how we read, learn, and communicate. The future promised by AI is written with stolen words.”
  • @dlberes Damon Beres on x
    Books3 is, to an extent, a known quantity, especially after a recent takedown request. (You'll see plenty of references in the piece.) This story represents the first comprehensive analysis of its contents for people outside of the AI community.
  • @srnlrsn @srnlrsn on x
    Update: AI is having a Napster moment. https://www.theatlantic.com/ ... [image]
  • @leifweatherby @leifweatherby on x
    amazing piece Alex and @dlberes - super important observations late in the article about the balance between copyright and distribution. if Books3 is taken out of LlaMA that will be a boon to OpenAI, at least for the moment. data wars are fully underway https://www.theatlantic.co…
  • @iamrobotbear @iamrobotbear on x
    @dlberes ... This is misleading as hell. Meta didn't release their LLaMA dataset first off. Secondly, look at the Google Books case and pls explain how this isn't transformative.
  • @garymarcus Gary Marcus on x
    “The future promised by AI is written with stolen words.” Literally:
  • @marc__watkins Marc Watkins on x
    To emphasize how FUBAR this whole situation is with training data that may have been pirated: [image]
  • @lmatsakis Louise Matsakis on x
    It's really instructive how the dataset here was analyzed. To understand how generative AI was built, researchers and journalists are going to need to rely on similar techniques https://www.theatlantic.com/ ... [image]
  • @heidilegg Heidi Legg on x
    First journalists, now authors. How Big Tech Platforms wiped out many essential workers by refusing to pay for our work.
  • @lmatsakis Louise Matsakis on x
    Meta trained LLaMA on upwards of *170,000* pirated books, the majority of which were written in the last 20 years. Great scoop from @TheAtlantic https://www.theatlantic.com/ ...
  • @mkirschenbaum Matthew Kirschenbaum on x
    This is a major, major piece of reporting and computer forensics at the intersect of #criticalAI, #bookhistory, and #dh. Hats off to @dlberes and of course the author, Alex Reisner. [image]
  • @stevesi Steven Sinofsky on x
    Generative-AI models by Meta and others have been trained on a secret trove of more than 170,000 books to gain their power. In a new analysis, Alex Reisner uncovers the authors whose work has been fed into popular AI programs: https://www.theatlantic.com/ ...
  • @kbandersen Kurt Andersen on x
    A corpus called Book3 used to train AIs by @Meta, @Bloomberg etc includes 170K+ post -2000 books, 1/3 fiction, 2/3 nonfiction. Thx for the piece, @TheAtlantic, but Alex Reisner is a programmer so please have them create an EZ search page for authors to see if our books were used.
  • @theatlantic @theatlantic on x
    Generative-AI models by Meta and others have been trained on a secret trove of more than 170,000 books to gain their power. In a new analysis, Alex Reisner uncovers the authors whose work has been fed into popular AI programs: https://www.theatlantic.com/ ...
  • @sam_fentress Sam Fentress on x
    you heard it here first: meta's LLM was trained on, among other things, 600 Verso titles https://www.theatlantic.com/ ...
  • @kz_howell K. Z. Howell on x
    As expected, plagiarism writ large. Artificial intelligence is neither artificial nor intelligent, it is theft on a scale even government is incapable of. https://www.theatlantic.com/ ...
  • @dlberes Damon Beres on x
    NEW: Meta, Bloomberg, and EleutherAI have trained generative AI on a dataset including upwards of 170,000 pirated books from authors like Stephen King, Zadie Smith, Margaret Atwood. Legality is complex. We have new details and context. tip @Techmeme https://www.theatlantic.com/ .…
  • @kylebarr5 Kyle Barr on x
    My latest for @Gizmodo hits on the cross-pollination of piracy ethics, copyright, and AI. Major companies like @MetaAI have trained their models on copyrighted works, but while tech giants can weather the storm of IP lawsuits and takedowns, small fry have a much harder time.
  • @lincodega @lincodega on x
    Copyright is always a difficult part of the law to navigate, but adding AI training is what will give a lot of power back to authors and IP owners.@KyleBarr5 nails it here. [image]
  • @theshawwn Shawn Presser on x
    A thoughtful Gizmodo article on books3 by @KyleBarr5: https://gizmodo.com/... Who would've guessed that the academictorrents website would be the last safe haven for AI research?
  • @jeffreygoldberg Jeffrey Goldberg on x
    A story about book pirates and AI: https://www.theatlantic.com/ ...
  • @ivanthek @ivanthek on x
    “The future promised by AI is written with stolen words.” The Achilles Heel of the current LLM craze. Lawyers will get paid. https://www.theatlantic.com/ ...
  • r/technology r on reddit
    Revealed: The Authors Whose Pirated Books Are Powering Generative AI |  Stephen King, Zadie Smith, and Michael Pollan are among thousands …