Penguin Random House amends its copyright notice globally to prohibit the use of books for training AI; the notice will be included in new titles and reprints
I'd love them to go ahead and sue any end user of generative AI whose content appears to contain their copyrighted material. … X: Paris Marx / @parismarx : “No part of this book may be used or reproduced in any manner for the purpose of training artificial intelligence technologies or systems.” good move from penguin https://www.theverge.com/... Trevor Baylis / @trevylimited : Note this mentions “the purpose of training artificial intelligence technologies or systems” - And not “Text and Data Mining.” Publisher's lawyers are understanding the difference. There are no © exceptions for Machine Learning. LinkedIn: Suzanne Arnold : Very interested to see Penguin Random House adding “No part of this book may be used or reproduced in any manner for the purpose … Forums: Hacker News : Penguin Random House underscores copyright protection in AI rebuff Msmash / Slashdot : Penguin Random House Underscores Copyright Protection in AI Rebuff See also Mediagazer
Context & Ripple Effects
Penguin Random House’s notice formalizes a publisher-level boundary around books as model-training inputs, amid efforts to remove the Books3 dataset from circulation and ongoing uncertainty over what copyright and fair-use doctrine permits. The move creates a clear contractual and evidentiary signal even though a notice alone does not resolve those legal questions.
The coverage later splits into two paths: a ruling that distinguished training from retaining pirated copies in a case involving Anthropic’s book corpus, and Johns Hopkins University Press’s decision to license authors’ books for AI training. Together, they show publishers testing both restriction and licensing as routes to control.
First-order effects
- New Penguin Random House titles and reprints will carry an explicit prohibition on AI-training use, putting model developers and data suppliers on notice of the publisher’s stated terms.
- The publisher gains a more consistent record of objection across its catalog, potentially strengthening its position in future licensing discussions or disputes over newly acquired copies.
Second-order effects
- AI developers and dataset intermediaries face greater provenance and permissions scrutiny for books acquired after the notice is deployed; a blanket assumption that purchased or accessible text is usable becomes harder to defend.
- Other publishers can adopt similar language or use it as leverage for paid access, while the alternative route—direct licensing of books for training—becomes more salient for developers seeking lower-risk supply.
Third-order effects
- If publisher notices become standard, book training data may shift from broadly collected corpora toward traceable, licensed or explicitly authorized catalogs, although the enforceability of notices will still depend on courts and applicable law.
- The market is moving toward separating access to a work from permission to use it as training input, a boundary central to copyright debates around AI training and model development.
The trend: Publishers are treating AI training rights as a distinct commercial and legal layer, pairing opt-outs with selective licensing as the rules for training-data access remain contested.