/
Navigation
Chronicles
Browse all articles
Explore
Semantic exploration
Research
Entity momentum
Nexus
Correlations & relationships
Story Arc
Topic evolution
Drift Map
Semantic trajectory animation
Posts
Analysis & commentary
Pulse API
Tech news intelligence API
Browse
Entities
Companies, people, products, technologies
Domains
Browse by publication source
Handles
Browse by social media handle
Detection
Concept Search
Semantic similarity search
High Impact Stories
Top coverage by position
Sentiment Analysis
Positive/negative coverage
Anomaly Detection
Unusual coverage patterns
Analysis
Rivalry Report
Compare two entities head-to-head
Semantic Pivots
Narrative discontinuities
Crisis Response
Event recovery patterns
Connected
Search: /
Command: ⌘K
Embeddings: large
TEXXR

Chronicles

The story behind the story

days · browse · Enter similar · o open

Copyright activists are working to wipe Books3 from the internet, which may only benefit the big companies that have already been using the AI training dataset

Because people use the term Artificial Intelligence to cover LLMs and other generative tools, they confuse the category of “people” with the category of “tools”.  —  I am in no way denying that someday, a person might be replicated or born in software. … X: @knibbs : My latest dispatch from the AI backlash beat, featuring Sarah Silverman's lawyer, Meta, Aaron Swartz's code, a chatty midwestern open-access activist, and around 197,000 books: https://www.wired.com/... Chris Harihar / @chrisharihar : @Knibbs Good article! Not sure of the solution, but legislation removing book text datasets from the open web so AIs can't train on them feels like using gum on a sinking boat. Yeah, it might hold for a bit, but the water's still coming in. LinkedIn: Eryk Salvaggio : Was asked for a comment on datasets and consent for this story in Wired Magazine.  What it comes down to is that nobody wants to go find 100,000 public domain images … Forums: Hacker News : The Battle over Books3

Wired Kate Knibbs

Context & Ripple Effects

The dispute follows a broader fight over who controls digitized books: publishers had already challenged Internet Archive’s ebook practices in a case centered on ownership of digital copies. It also arrives shortly after writer backlash shuttered Prosecraft, a book-analysis site whose corpus drew concern over possible AI use after authors objected to its dataset.

The key tension is distributional rather than simply legal: restricting public access to a corpus can affect future model builders differently from companies that have already had access to it.

First-order effects

  • People hosting, mirroring, or relying on Books3 face pressure to remove or avoid the dataset, reducing straightforward access for researchers and smaller AI developers.
  • Companies that have already used Books3 are comparatively insulated from a takedown effort; the reported action does not by itself reverse prior training use.

Second-order effects

  • The gap in usable training material could push newer entrants toward licensed, proprietary, or internally assembled text collections, raising the importance of rights-clearance capabilities.
  • Authors and publishers gain a clearer bargaining lever as the availability of large book corpora becomes contested, reinforcing calls for collective compensation negotiations over creators’ share of AI value.

Third-order effects

  • If removals of openly circulated corpora become common, training-data access may become a durable competitive moat: incumbents with prior access or licensing budgets would hold an advantage over later entrants.
  • The episode points toward a more formal permission boundary for public and digitized text, though courts and policy—not takedowns alone—will determine how broadly that boundary applies.

The trend: Generative AI is moving from an era of freely assembled web-scale corpora toward a market where provenance, permission, and prior data access shape competitive power.

Discussion

  • @knibbs @knibbs on x
    My latest dispatch from the AI backlash beat, featuring Sarah Silverman's lawyer, Meta, Aaron Swartz's code, a chatty midwestern open-access activist, and around 197,000 books: https://www.wired.com/...
  • @chrisharihar Chris Harihar on x
    @Knibbs Good article! Not sure of the solution, but legislation removing book text datasets from the open web so AIs can't train on them feels like using gum on a sinking boat. Yeah, it might hold for a bit, but the water's still coming in.