/
Navigation
Chronicles
Browse all articles
Explore
Semantic exploration
Research
Entity momentum
Nexus
Correlations & relationships
Story Arc
Topic evolution
Drift Map
Semantic trajectory animation
Posts
Analysis & commentary
Pulse API
Tech news intelligence API
Browse
Entities
Companies, people, products, technologies
Domains
Browse by publication source
Handles
Browse by social media handle
Detection
Concept Search
Semantic similarity search
High Impact Stories
Top coverage by position
Sentiment Analysis
Positive/negative coverage
Anomaly Detection
Unusual coverage patterns
Analysis
Rivalry Report
Compare two entities head-to-head
Semantic Pivots
Narrative discontinuities
Crisis Response
Event recovery patterns
Connected
Search: /
Command: ⌘K
Embeddings: large
TEXXR

Chronicles

The story behind the story

days · browse · Enter similar · o open

Penguin Random House amends its copyright notice globally to prohibit the use of books for training AI; the notice will be included in new titles and reprints

I'd love them to go ahead and sue any end user of generative AI whose content appears to contain their copyrighted material. … X: Paris Marx / @parismarx : “No part of this book may be used or reproduced in any manner for the purpose of training artificial intelligence technologies or systems.” good move from penguin https://www.theverge.com/... Trevor Baylis / @trevylimited : Note this mentions “the purpose of training artificial intelligence technologies or systems” - And not “Text and Data Mining.” Publisher's lawyers are understanding the difference. There are no © exceptions for Machine Learning. LinkedIn: Suzanne Arnold : Very interested to see Penguin Random House adding “No part of this book may be used or reproduced in any manner for the purpose … Forums: Hacker News : Penguin Random House underscores copyright protection in AI rebuff Msmash / Slashdot : Penguin Random House Underscores Copyright Protection in AI Rebuff See also Mediagazer

The Bookseller Matilda Battersby

Context & Ripple Effects

Penguin Random House’s notice formalizes a publisher-level boundary around books as model-training inputs, amid efforts to remove the Books3 dataset from circulation and ongoing uncertainty over what copyright and fair-use doctrine permits. The move creates a clear contractual and evidentiary signal even though a notice alone does not resolve those legal questions.

The coverage later splits into two paths: a ruling that distinguished training from retaining pirated copies in a case involving Anthropic’s book corpus, and Johns Hopkins University Press’s decision to license authors’ books for AI training. Together, they show publishers testing both restriction and licensing as routes to control.

First-order effects

  • New Penguin Random House titles and reprints will carry an explicit prohibition on AI-training use, putting model developers and data suppliers on notice of the publisher’s stated terms.
  • The publisher gains a more consistent record of objection across its catalog, potentially strengthening its position in future licensing discussions or disputes over newly acquired copies.

Second-order effects

  • AI developers and dataset intermediaries face greater provenance and permissions scrutiny for books acquired after the notice is deployed; a blanket assumption that purchased or accessible text is usable becomes harder to defend.
  • Other publishers can adopt similar language or use it as leverage for paid access, while the alternative route—direct licensing of books for training—becomes more salient for developers seeking lower-risk supply.

Third-order effects

  • If publisher notices become standard, book training data may shift from broadly collected corpora toward traceable, licensed or explicitly authorized catalogs, although the enforceability of notices will still depend on courts and applicable law.
  • The market is moving toward separating access to a work from permission to use it as training input, a boundary central to copyright debates around AI training and model development.

The trend: Publishers are treating AI training rights as a distinct commercial and legal layer, pairing opt-outs with selective licensing as the rules for training-data access remain contested.

Discussion

  • @parismarx Paris Marx on x
    “No part of this book may be used or reproduced in any manner for the purpose of training artificial intelligence technologies or systems.” good move from penguin https://www.theverge.com/...
  • @trevylimited Trevor Baylis on x
    Note this mentions “the purpose of training artificial intelligence technologies or systems” - And not “Text and Data Mining.” Publisher's lawyers are understanding the difference. There are no © exceptions for Machine Learning.