/
Navigation
Chronicles
Browse all articles
Explore
Semantic exploration
Research
Entity momentum
Nexus
Correlations & relationships
Story Arc
Topic evolution
Drift Map
Semantic trajectory animation
Posts
Analysis & commentary
Pulse API
Tech news intelligence API
Browse
Entities
Companies, people, products, technologies
Domains
Browse by publication source
Handles
Browse by social media handle
Detection
Concept Search
Semantic similarity search
High Impact Stories
Top coverage by position
Sentiment Analysis
Positive/negative coverage
Anomaly Detection
Unusual coverage patterns
Analysis
Rivalry Report
Compare two entities head-to-head
Semantic Pivots
Narrative discontinuities
Crisis Response
Event recovery patterns
Connected
Search: /
Command: ⌘K
Embeddings: large
TEXXR

Chronicles

The story behind the story

days · browse · Enter similar · o open

Kadrey v. Meta: unsealed emails show Meta allegedly torrented 81.7TB+ of data across multiple shadow libraries through the site Anna's Archive, for AI training

Newly unsealed emails allegedly provide the “most damning evidence” yet against Meta in a copyright case raised by book authors alleging …

Ars Technica Ashley Belanger

Context & Ripple Effects

The authors’ case had already been tied to allegations that Meta used LibGen material for Llama training, including internal-chat evidence that it sought to conceal that use. The newly unsealed correspondence adds a more detailed alleged acquisition trail to that dispute.

It also sharpens the contrast with Meta’s halted book-licensing outreach, which reportedly ran into slow uptake and logistical obstacles. The case later became a major early test of authors’ claims over model-training inputs.

First-order effects

  • The unsealed emails give the plaintiffs a more concrete alleged record to press in Kadrey v. Meta, while requiring Meta to contest the provenance, purpose, and legal significance of the claimed downloads.
  • The disclosure raises the stakes of discovery around Meta’s training-data controls and its decisions between licensed and shadow-library sources.

Second-order effects

  • Publishers and authors gain a more specific factual basis for demanding licensing, disclosure, or compensation from AI developers whose training-data practices are disputed.
  • Other model builders face added pressure to document dataset sourcing: alleged use of shadow libraries can become litigation evidence, not merely a copyright-theory debate.

Third-order effects

  • If courts treat acquisition records and internal deliberations as central evidence, AI copyright disputes may increasingly turn on auditable data provenance rather than broad arguments about training alone.
  • The pattern could favor formal licensing and traceable dataset governance, although the eventual effect depends on how courts resolve the underlying copyright claims.

The trend: AI copyright litigation is moving from abstract challenges to training toward evidence-driven scrutiny of how model developers obtained and governed their datasets.

Discussion

  • @jaypeg89 @jaypeg89 on bluesky
    Remember when Capitol Records sued a woman for downloading 24 songs on Kazaa and won a $1.92 million settlement?  How much is 80TB worth...  [embedded post]
  • @clearspark @clearspark on bluesky
    Torrent=file sharing.  Shadow libraries=pirate libraries of eBooks.  —  I randomly looked up the file size of eBooks and I got very different numbers.  So, let's pick a size of 50 MB.  —  So if I did this right that's like 1,713,373 copyright violations assuming an average book s…
  • @jackkennedy.ie Jack Kennedy on bluesky
    I have this funny feeling that Meta will not be pursued by prosecutors with the same murderous aggression faced by, say, Aaron Swartz [embedded post]
  • @ambaazaad Amba Azaad on bluesky
    “Meta also allegedly modified settings “so that the smallest amount of seeding possible could occur”  —  Meta is a leecher, literally.  [embedded post]
  • @ianbetteridge.com Ian Betteridge on bluesky
    I think under the computer misuse act this merits a maximum penalty of $1 million in fines and 35 years in prison for Zuck.  [embedded post]
  • @nixCraft@mastodon.social @nixCraft@mastodon.social on mastodon
    Newly unsealed emails allegedly provide the “most damning evidence” yet against Meta in a copyright case raised by book authors alleging that Meta illegally trained its AI models on pirated books.  Mark is truly doing everything to avoid paying authors, remove fact-checkers, and …
  • @bauerkahan Assemblymember Rebecca Bauer-Kahan on x
    Authors have a right to control and profit off of their own intellectual property. My bill, AB 412, requires developers to be transparent about their use of copyrighted materials. https://arstechnica.com/...
  • @porridgeawful @porridgeawful on x
    Hey look, it's that thing we all knew: “Torrenting from a corporate laptop doesn't feel right”: Meta emails unsealed - Ars Technica https://arstechnica.com/... Sure is open-source when you're hiding your sources, amirite? [image]
  • @jason_kint Jason Kint on x
    Oomph. Same playbook, over and over. capture as much data however possible to accelerate growth. zuckerberg wants to win sooooo badly. 1/2 [image]
  • r/NoShitSherlock r on reddit
    Meta torrented over 81.7TB of pirated books to train AI, authors say
  • r/neoliberal r on reddit
    “Torrenting from a corporate laptop doesn't feel right”: Meta emails unsealed
  • r/aiwars r on reddit
    Meta torrented over 81.7TB of pirated books to train AI, authors say
  • r/LinusTechTips r on reddit
    Meta torrented over 81.7TB of pirated books to train AI, authors say
  • r/books r on reddit
    Proof that Meta torrented “at least 81.7 terabytes of data” uncovered in a copyright case raised by book authors.
  • r/Annas_Archive r on reddit
    Meta torrented over 81.7TB of pirated books to train AI, authors say
  • r/libgen r on reddit
    Meta torrented over 81.7TB of pirated books to train AI, authors say
  • r/singularity r on reddit
    Meta torrented over 100 terabytes of pirated books from Anna's Archive, Z-Library and LibGen for AI training
  • r/trackers r on reddit
    Meta torrented 81.7 terabytes of data to train their AI models
  • r/Piracy r on reddit
    Meta Torrents Million of Books to Train Models
  • r/facebook r on reddit
    Likely the Largest Case of IP Theft through out the Entire History of Humanity
  • r/ArtistHate r on reddit
    There's still good news for humanity: Meta torrented every book ever written train AI, and their AI still sucks
  • r/technology r on reddit
    Meta torrented over 81.7TB of pirated books to train AI, authors say
  • r/technews r on reddit
    “Torrenting from a corporate laptop doesn't feel right”: Meta emails unsealed |  Meta's alleged torrenting and seeding of pirated books complicates copyright case.
  • r/artificial r on reddit
    Meta torrented over 81.7TB of pirated books to train AI, authors say
  • r/NewsWithJingjing r on reddit
    Meta torrented over 81.7TB of pirated books to train AI, authors say
  • r/economy r on reddit
    Meta torrented over 81.7TB of pirated books to train AI, authors say