/
Navigation
Chronicles
Browse all articles
Explore
Semantic exploration
Research
Entity momentum
Nexus
Correlations & relationships
Story Arc
Topic evolution
Drift Map
Semantic trajectory animation
Posts
Analysis & commentary
Pulse API
Tech news intelligence API
Browse
Entities
Companies, people, products, technologies
Domains
Browse by publication source
Handles
Browse by social media handle
Detection
Concept Search
Semantic similarity search
High Impact Stories
Top coverage by position
Sentiment Analysis
Positive/negative coverage
Anomaly Detection
Unusual coverage patterns
Analysis
Rivalry Report
Compare two entities head-to-head
Semantic Pivots
Narrative discontinuities
Crisis Response
Event recovery patterns
Connected
Search: /
Command: ⌘K
Embeddings: large
TEXXR

Chronicles

The story behind the story

days · browse · Enter similar · o open

A US judge rules Anthropic's use of copyrighted books to train AI was fair use, but its storage of pirated books in a central library used for training was not

Context: Last August, book authors Andrea Bartz, Charles Graeber and Kirk Wallace Johnson filed a copyright infringement class action …

ai fray Olivia Sophie Rafferty

Context & Ripple Effects

This ruling drew a line between model training and the acquisition-and-retention practices behind a training corpus. A contemporaneous Meta ruling on Llama book training likewise favored fair use, while leaving the quality of the plaintiffs’ case central to the outcome.

The dispute did not end with this decision: the authors later won permission to pursue a broader class action covering affected U.S. writers, and coverage subsequently reported preliminary approval of an Anthropic authors’ settlement. That sequence makes the sourcing of datasets, not just the act of training, a material legal exposure.

First-order effects

  • Anthropic receives support for the proposition that transformative use of books in model training can be fair use, but faces liability risk over its centralized store of pirated copies.
  • The author plaintiffs retain a viable path to challenge the allegedly unauthorized library, separating claims over corpus provenance from claims over model training.

Second-order effects

  • Other AI developers gain a useful litigation distinction but must document how training materials were obtained, stored, and governed; a fair-use defense for training does not resolve alleged piracy in the supply chain.
  • Publishers and authors have a narrower, more concrete target for claims and negotiations: earlier surviving claims over OpenAI’s use of books show that training-data disputes can proceed on theories beyond a single copyright argument.

Third-order effects

  • If courts continue to distinguish transformative training from unlawful corpus assembly, compliance will shift toward auditable data provenance, retention controls, and licensed or otherwise defensible content pipelines.
  • The result could produce a bifurcated market: model-training law may remain fact-specific, while the commercial value of well-governed training corpora rises as litigation risk concentrates upstream.

The trend: AI copyright disputes are moving from a broad question of whether training is permissible toward a more operational test of how training data is sourced, retained, and governed.

Discussion

  • @esqueer.net Alejandra Caraballo on bluesky
    Most of the AI companies used far more than lawfully acquired copies of books so this could present a major issue going forward.  We'll need to see how lawsuits involving scraped material off the internet play out to see how AI companies will fair under copyright law.
  • @esqueer.net Alejandra Caraballo on bluesky
    This presents a massive liability for AI companies still who often scraped millions of documents, videos, images etc. without authorization.  So the precedent this sets here in this lawsuit is that when lawfully acquired, training AI models on copyrighted material is fair use.
  • @reckless Nilay Patel on bluesky
    Alsup is a very smart and very sharp judge, and his order is eminently readable.  You should read it! www.documentcloud.org/documents/ 25...
  • @reckless Nilay Patel on bluesky
    If you are running away saying this case definitively rules the training an LLM is fair use... you're going to make some big and potentially very expensive mistakes.  Ruling very specifically does not reach *outputs* which feels important to future cases - and very clearly comes …
  • @jtlg James Grimmelmann on bluesky
    This decision will almost certainly be appealed, of course, but it may be a good bellwether for where these lawsuits are going in general.  If this pattern holds, then AI training will typically be fair use, but companies will need to turn square corners in acquiring their traini…
  • @jtlg James Grimmelmann on bluesky
    The big unanswered question (because it wasn't presented here) is whether web scraping is more like scanning books (fair use) or like downloading “pirated” books (not fair use).  —  (I put “pirated” in quotes because the distinction could come under pressure in future cases about…
  • @reckless Nilay Patel on bluesky
    Court rules Anthropic training an LLM on books is fair use... if they bought the books.  But Anthropic also pirated a lot of books, and now faces a second, potentially major, damages trial for stealing them [image]
  • @ballmatthew Matthew Ball on x
    U.S. District Court finds Anthropic's use of copyright books (i.e. w/o express consent/agreement/deal) to train LLMs is “transformative” ("among the most transformative in our lifetime") thus justified as fair use + “in light of the purposes of copyright” https://www.documentclou…
  • @eriqgardner Eriq Gardner on x
    🚨BREAKING: Federal judge concludes that using copyrighted works to train generative A.I. is transformative and ultimately a fair use. (Nevertheless, Anthropic can't beat the lawsuit because it pirated books for another purpose too.) First of kind ruling. https://www.documentcloud…
  • @neilturkewitz Neil Turkewitz on x
    Actually, that's not apparently what the court ruled. I haven't read the decision, but according to the article: “Alsup also said, however, that Anthropic's storage of the authors' books in a ‘central library’ violated their copyrights & was not fair use.” So not so fast!
  • @kimmonismus @kimmonismus on x
    This could be a landmark ruling: A court rules that anthropic models trained with licensed books fall under “fair use” and may be used. This is a major and significant ruling for the training of AI models - and at the same time, to be honest, a challenge for creative authors.
  • @matdryhurst Mat Dryhurst on x
    Big, and goes even further than I expected Judge determines the use of books for model training is transformative and constitutes fair use so long as outputs are not infringing Case is going to trial over the use of pirated books, which is obviously illegal and expected
  • @xlr8harder @xlr8harder on x
    Sounds like a huge win for fair use. Just should have bought the books first instead of stealing them.
  • @stevesi Steven Sinofsky on x
    Full ruling here. Expect future cases to spend more energy on substitute for the original. Not sure this is a final word or they made the best case. https://storage.courtlistener.com/ ...
  • @morqon Morgan on x
    anthropic wins a major judgement on fair use, can train models on purchased books, but goes to trial for storing pirated works - training is “spectacularly transformative” - memorisation of style and content for statistical modelling is analogous to a person learning from
  • @rainisto Roope Rainisto on x
    Really fascinating fair use analysis and summary judgment in favour of Anthropic (for nerds like me interested in the legal argument) - tens of pages of arguments about each of the four factors of fair use. [image]
  • r/singularity r on reddit
    A federal judge has ruled that Anthropic's use of books to train Claude falls under fair use, and is legal under U.S. copyright law
  • r/technology r on reddit
    Anthropic wins key ruling on AI in authors' copyright lawsuit