/
Navigation
Chronicles
Browse all articles
Explore
Semantic exploration
Research
Entity momentum
Nexus
Correlations & relationships
Story Arc
Topic evolution
Drift Map
Semantic trajectory animation
Posts
Analysis & commentary
Pulse API
Tech news intelligence API
Browse
Entities
Companies, people, products, technologies
Domains
Browse by publication source
Handles
Browse by social media handle
Detection
Concept Search
Semantic similarity search
High Impact Stories
Top coverage by position
Sentiment Analysis
Positive/negative coverage
Anomaly Detection
Unusual coverage patterns
Analysis
Rivalry Report
Compare two entities head-to-head
Semantic Pivots
Narrative discontinuities
Crisis Response
Event recovery patterns
Connected
Search: /
Command: ⌘K
Embeddings: large
TEXXR

Chronicles

The story behind the story

days · browse · Enter similar · o open

Kadrey v. Meta, centered on the use of LibGen to train Llama AI models, kicks off, marking the first big legal test in the ongoing battle over AI and copyright

Tech giant faces lawsuit from US authors over use of material from shadow library LibGen  —  Meta will fight a group of US authors …

Financial Times

Context & Ripple Effects

The trial follows a judge's decision to let the authors' claims proceed, while discovery had already put Meta's alleged use of LibGen and other shadow-library material under scrutiny through internal discussions about LibGen training data and unsealed allegations of large-scale shadow-library downloads.

It matters because the dispute focuses not only on whether books can be used in model training, but on the provenance of the training corpus. Meta had also reportedly paused book-licensing outreach after slow publisher uptake, sharpening the contrast between negotiated access and disputed data acquisition.

First-order effects

  • Meta must defend the sourcing and use of material behind Llama training against authors seeking to establish that those acts infringed their rights.
  • The authors gain a major forum to test whether alleged use of shadow-library copies changes the legal assessment of AI training.

Second-order effects

  • AI developers using broad web or library-derived corpora face greater pressure to document data provenance and distinguish licensed, public, and disputed sources.
  • Publishers and rights holders gain leverage in licensing discussions if litigation makes poorly documented training datasets a more material legal and reputational exposure.

Third-order effects

  • If courts treat corpus provenance as central rather than incidental, AI competition may shift toward auditable data supply chains and negotiated content access, favoring firms able to fund them.
  • The case is part of an unresolved boundary-setting process for copyright and model training; outcomes will determine whether licensing becomes a durable operating cost or a narrower remedy for particular sourcing practices.

The trend: Generative-AI developers are moving from scale-at-all-costs data collection toward a contest over whether training-data provenance can become a competitive and legal constraint.

Discussion

  • @cyrilpedia Thiago Carvalho on bluesky
    'The case, which has been brought by about a dozen authors including Ta-Nehisi Coates and Richard Kadrey, is centred around the $1.4tn social media giant's use of LibGen, a so-called shadow library of millions of books, academic articles and comics, to train its Llama AI models.'
  • @timoreilly Tim O'Reilly on x
    I haven't been following the various cases on AI and copyright, but ChatGPTiseatingtheworld does just that. Here's the fascinating analysis of the judge's questions in tomorrow's hearing on Kadrey v Meta, which, after disposing of all the other expected fair use issues, turns on