/
Navigation
Chronicles
Browse all articles
Explore
Semantic exploration
Research
Entity momentum
Nexus
Correlations & relationships
Story Arc
Topic evolution
Drift Map
Semantic trajectory animation
Posts
Analysis & commentary
Pulse API
Tech news intelligence API
Browse
Entities
Companies, people, products, technologies
Domains
Browse by publication source
Handles
Browse by social media handle
Detection
Concept Search
Semantic similarity search
High Impact Stories
Top coverage by position
Sentiment Analysis
Positive/negative coverage
Anomaly Detection
Unusual coverage patterns
Analysis
Rivalry Report
Compare two entities head-to-head
Semantic Pivots
Narrative discontinuities
Crisis Response
Event recovery patterns
Connected
Search: /
Command: ⌘K
Embeddings: large
TEXXR

Chronicles

The story behind the story

days · browse · Enter similar · o open

Analysis: Books3, a dataset used to train Meta's Llama, BloombergGPT, and EleutherAI's GPT-J, contains 170K+ books from Stephen King and other authors

One of the most troubling issues around generative AI is simple: It's being made in secret.  To produce humanlike answers to questions …

The Atlantic Alex Reisner

Context & Ripple Effects

The Books3 disclosure put named commercial and open-model projects in the same training-data controversy, making the provenance of literary corpora a concrete issue rather than an abstract concern. Follow-on coverage showed copyright activists seeking to remove the dataset from circulation through a campaign to wipe Books3 from the internet.

The issue fits a broader pipeline: later reporting examined LibGen as a major pirate-library source allegedly used for AI training, while researchers later found that leading models could produce long book excerpts under strategic prompting. Together, those developments sharpen the distinction between access to text for training and control over how models may reproduce it.

First-order effects

  • Authors and rightsholders whose books appear in Books3 gain a clearer factual basis to scrutinize the training practices behind Meta's Llama, BloombergGPT, and GPT-J.
  • The named model developers face immediate provenance and reputational pressure because a dataset containing more than 170,000 books is publicly tied to their systems.

Second-order effects

  • Efforts to remove Books3 may reduce easy access for later entrants, while companies that already trained on it retain the resulting model capabilities—an asymmetry highlighted by the push to take Books3 offline.
  • Publishers and AI developers have stronger incentives to distinguish licensed or documented text sources from corpora associated with pirate-library distribution, including the LibGen training-data allegations.

Third-order effects

  • If disclosures and reproducibility findings continue to accumulate, training-data provenance is likely to become a competitive and governance constraint alongside model quality, especially for systems built on cultural works.
  • The industry may move toward more explicit licensing, auditing, and output safeguards; whether that narrows the gap between established model owners and smaller developers depends on how access rules are applied.

The trend: Generative AI is shifting from opaque web-scale data collection toward a contested market for traceable, governable rights to high-value training content.

Discussion

  • @technursejon@mastodon.art Jon on mastodon
    “The exploitation of pirated books for profit, with the goal of replacing the writers whose work was taken—this is a different and disturbing trend.”  #AI #ArtificialIntelligence  —  https://www.theatlantic.com/ ...
  • @mkirschenbaum Matthew Kirschenbaum on x
    This is a major, major piece of reporting and computer forensics at the intersect of #criticalAI, #bookhistory, and #dh. Hats off to @dlberes and of course the author, Alex Reisner. [image]
  • @kbandersen Kurt Andersen on x
    A corpus called Book3 used to train AIs by @Meta, @Bloomberg etc includes 170K+ post -2000 books, 1/3 fiction, 2/3 nonfiction. Thx for the piece, @TheAtlantic, but Alex Reisner is a programmer so please have them create an EZ search page for authors to see if our books were used.
  • @stevesi Steven Sinofsky on x
    Generative-AI models by Meta and others have been trained on a secret trove of more than 170,000 books to gain their power. In a new analysis, Alex Reisner uncovers the authors whose work has been fed into popular AI programs: https://www.theatlantic.com/ ...
  • @dlberes Damon Beres on x
    NEW: Meta, Bloomberg, and EleutherAI have trained generative AI on a dataset including upwards of 170,000 pirated books from authors like Stephen King, Zadie Smith, Margaret Atwood. Legality is complex. We have new details and context. tip @Techmeme https://www.theatlantic.com/ .…
  • @theshawwn Shawn Presser on x
    A thoughtful Gizmodo article on books3 by @KyleBarr5: https://gizmodo.com/... Who would've guessed that the academictorrents website would be the last safe haven for AI research?
  • @theatlantic @theatlantic on x
    Generative-AI models by Meta and others have been trained on a secret trove of more than 170,000 books to gain their power. In a new analysis, Alex Reisner uncovers the authors whose work has been fed into popular AI programs: https://www.theatlantic.com/ ...
  • @sam_fentress Sam Fentress on x
    you heard it here first: meta's LLM was trained on, among other things, 600 Verso titles https://www.theatlantic.com/ ...
  • @kylebarr5 Kyle Barr on x
    My latest for @Gizmodo hits on the cross-pollination of piracy ethics, copyright, and AI. Major companies like @MetaAI have trained their models on copyrighted works, but while tech giants can weather the storm of IP lawsuits and takedowns, small fry have a much harder time.
  • @jeffreygoldberg Jeffrey Goldberg on x
    A story about book pirates and AI: https://www.theatlantic.com/ ...
  • @lincodega @lincodega on x
    Copyright is always a difficult part of the law to navigate, but adding AI training is what will give a lot of power back to authors and IP owners.@KyleBarr5 nails it here. [image]
  • @garymarcus Gary Marcus on x
    “The future promised by AI is written with stolen words.” Literally:
  • @ivanthek @ivanthek on x
    “The future promised by AI is written with stolen words.” The Achilles Heel of the current LLM craze. Lawyers will get paid. https://www.theatlantic.com/ ...
  • @kz_howell K. Z. Howell on x
    As expected, plagiarism writ large. Artificial intelligence is neither artificial nor intelligent, it is theft on a scale even government is incapable of. https://www.theatlantic.com/ ...