Kadrey v. Meta: unsealed emails show Meta allegedly torrented 81.7TB+ of data across multiple shadow libraries through the site Anna's Archive, for AI training
Newly unsealed emails allegedly provide the “most damning evidence” yet against Meta in a copyright case raised by book authors alleging …
Ars TechnicaAshley Belanger
Context & Ripple Effects
The authors’ case had already been tied to allegations that Meta used LibGen material for Llama training, including internal-chat evidence that it sought to conceal that use. The newly unsealed correspondence adds a more detailed alleged acquisition trail to that dispute.
It also sharpens the contrast with Meta’s halted book-licensing outreach, which reportedly ran into slow uptake and logistical obstacles. The case later became a major early test of authors’ claims over model-training inputs.
First-order effects
The unsealed emails give the plaintiffs a more concrete alleged record to press in Kadrey v. Meta, while requiring Meta to contest the provenance, purpose, and legal significance of the claimed downloads.
The disclosure raises the stakes of discovery around Meta’s training-data controls and its decisions between licensed and shadow-library sources.
Second-order effects
Publishers and authors gain a more specific factual basis for demanding licensing, disclosure, or compensation from AI developers whose training-data practices are disputed.
Other model builders face added pressure to document dataset sourcing: alleged use of shadow libraries can become litigation evidence, not merely a copyright-theory debate.
Third-order effects
If courts treat acquisition records and internal deliberations as central evidence, AI copyright disputes may increasingly turn on auditable data provenance rather than broad arguments about training alone.
The pattern could favor formal licensing and traceable dataset governance, although the eventual effect depends on how courts resolve the underlying copyright claims.
The trend: AI copyright litigation is moving from abstract challenges to training toward evidence-driven scrutiny of how model developers obtained and governed their datasets.
Remember when Capitol Records sued a woman for downloading 24 songs on Kazaa and won a $1.92 million settlement? How much is 80TB worth... [embedded post]
Torrent=file sharing. Shadow libraries=pirate libraries of eBooks. — I randomly looked up the file size of eBooks and I got very different numbers. So, let's pick a size of 50 MB. — So if I did this right that's like 1,713,373 copyright violations assuming an average book s…
Newly unsealed emails allegedly provide the “most damning evidence” yet against Meta in a copyright case raised by book authors alleging that Meta illegally trained its AI models on pirated books. Mark is truly doing everything to avoid paying authors, remove fact-checkers, and …
@bauerkahan Assemblymember Rebecca Bauer-Kahan on x
Authors have a right to control and profit off of their own intellectual property. My bill, AB 412, requires developers to be transparent about their use of copyrighted materials. https://arstechnica.com/...
Hey look, it's that thing we all knew: “Torrenting from a corporate laptop doesn't feel right”: Meta emails unsealed - Ars Technica https://arstechnica.com/... Sure is open-source when you're hiding your sources, amirite? [image]
“Torrenting from a corporate laptop doesn't feel right”: Meta emails unsealed | Meta's alleged torrenting and seeding of pirated books complicates copyright case.