A US judge rules Anthropic's use of copyrighted books to train AI was fair use, but its storage of pirated books in a central library used for training was not
This ruling drew a line between model training and the acquisition-and-retention practices behind a training corpus. A contemporaneous Meta ruling on Llama book training likewise favored fair use, while leaving the quality of the plaintiffs’ case central to the outcome.
The dispute did not end with this decision: the authors later won permission to pursue a broader class action covering affected U.S. writers, and coverage subsequently reported preliminary approval of an Anthropic authors’ settlement. That sequence makes the sourcing of datasets, not just the act of training, a material legal exposure.
First-order effects
Anthropic receives support for the proposition that transformative use of books in model training can be fair use, but faces liability risk over its centralized store of pirated copies.
The author plaintiffs retain a viable path to challenge the allegedly unauthorized library, separating claims over corpus provenance from claims over model training.
Second-order effects
Other AI developers gain a useful litigation distinction but must document how training materials were obtained, stored, and governed; a fair-use defense for training does not resolve alleged piracy in the supply chain.
Publishers and authors have a narrower, more concrete target for claims and negotiations: earlier surviving claims over OpenAI’s use of books show that training-data disputes can proceed on theories beyond a single copyright argument.
Third-order effects
If courts continue to distinguish transformative training from unlawful corpus assembly, compliance will shift toward auditable data provenance, retention controls, and licensed or otherwise defensible content pipelines.
The result could produce a bifurcated market: model-training law may remain fact-specific, while the commercial value of well-governed training corpora rises as litigation risk concentrates upstream.
The trend:AI copyright disputes are moving from a broad question of whether training is permissible toward a more operational test of how training data is sourced, retained, and governed.
Most of the AI companies used far more than lawfully acquired copies of books so this could present a major issue going forward. We'll need to see how lawsuits involving scraped material off the internet play out to see how AI companies will fair under copyright law.
This presents a massive liability for AI companies still who often scraped millions of documents, videos, images etc. without authorization. So the precedent this sets here in this lawsuit is that when lawfully acquired, training AI models on copyrighted material is fair use.
If you are running away saying this case definitively rules the training an LLM is fair use... you're going to make some big and potentially very expensive mistakes. Ruling very specifically does not reach *outputs* which feels important to future cases - and very clearly comes …
This decision will almost certainly be appealed, of course, but it may be a good bellwether for where these lawsuits are going in general. If this pattern holds, then AI training will typically be fair use, but companies will need to turn square corners in acquiring their traini…
The big unanswered question (because it wasn't presented here) is whether web scraping is more like scanning books (fair use) or like downloading “pirated” books (not fair use). — (I put “pirated” in quotes because the distinction could come under pressure in future cases about…
Court rules Anthropic training an LLM on books is fair use... if they bought the books. But Anthropic also pirated a lot of books, and now faces a second, potentially major, damages trial for stealing them [image]
U.S. District Court finds Anthropic's use of copyright books (i.e. w/o express consent/agreement/deal) to train LLMs is “transformative” ("among the most transformative in our lifetime") thus justified as fair use + “in light of the purposes of copyright” https://www.documentclou…
🚨BREAKING: Federal judge concludes that using copyrighted works to train generative A.I. is transformative and ultimately a fair use. (Nevertheless, Anthropic can't beat the lawsuit because it pirated books for another purpose too.) First of kind ruling. https://www.documentcloud…
Actually, that's not apparently what the court ruled. I haven't read the decision, but according to the article: “Alsup also said, however, that Anthropic's storage of the authors' books in a ‘central library’ violated their copyrights & was not fair use.” So not so fast!
This could be a landmark ruling: A court rules that anthropic models trained with licensed books fall under “fair use” and may be used. This is a major and significant ruling for the training of AI models - and at the same time, to be honest, a challenge for creative authors.
Big, and goes even further than I expected Judge determines the use of books for model training is transformative and constitutes fair use so long as outputs are not infringing Case is going to trial over the use of pirated books, which is obviously illegal and expected
Full ruling here. Expect future cases to spend more energy on substitute for the original. Not sure this is a final word or they made the best case. https://storage.courtlistener.com/ ...
anthropic wins a major judgement on fair use, can train models on purchased books, but goes to trial for storing pirated works - training is “spectacularly transformative” - memorisation of style and content for statistical modelling is analogous to a person learning from
Really fascinating fair use analysis and summary judgment in favour of Anthropic (for nerds like me interested in the legal argument) - tens of pages of arguments about each of the four factors of fair use. [image]