Copyright activists are working to wipe Books3 from the internet, which may only benefit the big companies that have already been using the AI training dataset
Because people use the term Artificial Intelligence to cover LLMs and other generative tools, they confuse the category of “people” with the category of “tools”. — I am in no way denying that someday, a person might be replicated or born in software. … X: @knibbs : My latest dispatch from the AI backlash beat, featuring Sarah Silverman's lawyer, Meta, Aaron Swartz's code, a chatty midwestern open-access activist, and around 197,000 books: https://www.wired.com/... Chris Harihar / @chrisharihar : @Knibbs Good article! Not sure of the solution, but legislation removing book text datasets from the open web so AIs can't train on them feels like using gum on a sinking boat. Yeah, it might hold for a bit, but the water's still coming in. LinkedIn: Eryk Salvaggio : Was asked for a comment on datasets and consent for this story in Wired Magazine. What it comes down to is that nobody wants to go find 100,000 public domain images … Forums: Hacker News : The Battle over Books3
Context & Ripple Effects
The dispute follows a broader fight over who controls digitized books: publishers had already challenged Internet Archive’s ebook practices in a case centered on ownership of digital copies. It also arrives shortly after writer backlash shuttered Prosecraft, a book-analysis site whose corpus drew concern over possible AI use after authors objected to its dataset.
The key tension is distributional rather than simply legal: restricting public access to a corpus can affect future model builders differently from companies that have already had access to it.
First-order effects
- People hosting, mirroring, or relying on Books3 face pressure to remove or avoid the dataset, reducing straightforward access for researchers and smaller AI developers.
- Companies that have already used Books3 are comparatively insulated from a takedown effort; the reported action does not by itself reverse prior training use.
Second-order effects
- The gap in usable training material could push newer entrants toward licensed, proprietary, or internally assembled text collections, raising the importance of rights-clearance capabilities.
- Authors and publishers gain a clearer bargaining lever as the availability of large book corpora becomes contested, reinforcing calls for collective compensation negotiations over creators’ share of AI value.
Third-order effects
- If removals of openly circulated corpora become common, training-data access may become a durable competitive moat: incumbents with prior access or licensing budgets would hold an advantage over later entrants.
- The episode points toward a more formal permission boundary for public and digitized text, though courts and policy—not takedowns alone—will determine how broadly that boundary applies.
The trend: Generative AI is moving from an era of freely assembled web-scale corpora toward a market where provenance, permission, and prior data access shape competitive power.