Kadrey v. Meta, centered on the use of LibGen to train Llama AI models, kicks off, marking the first big legal test in the ongoing battle over AI and copyright
Tech giant faces lawsuit from US authors over use of material from shadow library LibGen — Meta will fight a group of US authors …
Context & Ripple Effects
The trial follows a judge's decision to let the authors' claims proceed, while discovery had already put Meta's alleged use of LibGen and other shadow-library material under scrutiny through internal discussions about LibGen training data and unsealed allegations of large-scale shadow-library downloads.
It matters because the dispute focuses not only on whether books can be used in model training, but on the provenance of the training corpus. Meta had also reportedly paused book-licensing outreach after slow publisher uptake, sharpening the contrast between negotiated access and disputed data acquisition.
First-order effects
- Meta must defend the sourcing and use of material behind Llama training against authors seeking to establish that those acts infringed their rights.
- The authors gain a major forum to test whether alleged use of shadow-library copies changes the legal assessment of AI training.
Second-order effects
- AI developers using broad web or library-derived corpora face greater pressure to document data provenance and distinguish licensed, public, and disputed sources.
- Publishers and rights holders gain leverage in licensing discussions if litigation makes poorly documented training datasets a more material legal and reputational exposure.
Third-order effects
- If courts treat corpus provenance as central rather than incidental, AI competition may shift toward auditable data supply chains and negotiated content access, favoring firms able to fund them.
- The case is part of an unresolved boundary-setting process for copyright and model training; outcomes will determine whether licensing becomes a durable operating cost or a narrower remedy for particular sourcing practices.
The trend: Generative-AI developers are moving from scale-at-all-costs data collection toward a contest over whether training-data provenance can become a competitive and legal constraint.