Court docs: Mark Zuckerberg approved the Meta team that trains AI Llama models to use data from LibGen, a “links aggregator” to pirated, copyrighted material
Counsel for plaintiffs in a copyright lawsuit filed against Meta allege that Meta CEO Mark Zuckerberg gave the green light …
TechCrunchKyle Wiggers
Context & Ripple Effects
The filing puts executive approval at the center of the Kadrey v. Meta dispute, which a federal judge later allowed to proceed. That makes the provenance of Llama’s training corpus a governance question, not only a technical one.
Related coverage subsequently pointed to internal discussions about LibGen use and concealment, while a later ruling found Meta’s book-training use fair use on the arguments presented. The sequence leaves the underlying boundary around AI training data contested rather than settled.
First-order effects
The allegation raises Meta’s litigation and reputational exposure by tying the claimed LibGen decision to Zuckerberg’s approval, rather than treating it solely as a team-level sourcing choice.
Authors suing Meta gain a more specific account of how the disputed Llama training material was allegedly authorized, though the filing itself does not establish liability.
Second-order effects
Meta and other model developers face stronger incentives to document dataset provenance and decision-making, especially where material originates from sources associated with infringement.
The dispute increases the value of licensed or clearly permissioned corpora for developers seeking to reduce litigation uncertainty, even as courts assess fair-use arguments.
Third-order effects
If courts continue examining both training inputs and internal approval processes, AI-data governance may become a durable competitive and compliance function rather than an after-the-fact legal review.
The later fair-use ruling in Meta’s books case illustrates that outcomes can turn on the record and legal arguments, leaving no simple industry-wide clearance rule for copyrighted training data.
The trend: Generative-AI competition is pushing training-data provenance and permission boundaries into a central legal and corporate-governance battleground.
Meta tried to hide that it trained its AI on pirated books. — One staff member wrote: — “If there is media coverage suggesting we have used a dataset we know to be pirated, such as LibGen, this may undermine our negotiating position with regulators on these issues” — Fantas…
According to court documents, Mark Zuckerberg gave Meta's Llama team the OK to train on copyrighted works — techcrunch.com/2025/01/09/m... storage.courtlistener.com/recap/ gov.us... Ultimately this is going to impact on those deploying these models, including those trained us…
Well, it seems as though Meta pirated all our books and wrote special code to strip them of copyright and acknowledgments so they wouldn't get caught out. techcrunch.com/2025/01/09/m...
“If there is media coverage suggesting we have used a dataset we know to be pirated, such as LibGen, this may undermine our negotiating position with regulators on these issues.”
As a new Twitch streamer I've taken pains to avoid violating copyright law with the music I use, the images I show and anything else that appears on-stream. — Meanwhile tech giants allegedly casually use copyrighted material to train AI with barely a blink of an eye techcrunch.…
So this is personal opinion, I'm not a copyright attorney. — But what Meta, Co-Pilot, Chat GPT et al are trying to do is amass so much market share, business integration etc that by the time the Supreme Court gets a case, enforcing copyright law is no longer possible. — techc…
“The filing quotes Meta employees as referring to LibGen as a “data set we know to be pirated,” and flagging that its use “may undermine [Meta's] negotiating position with regulators.
Not only did Zuck sign off on using copyrighted works, Meta “wrote a script to remove copyright info, including the word ‘copyright’ and ‘acknowledgments,’ from e-books in LibGen...” — ... Meta actively tried to conceal it. — techcrunch.com/2025/01/09/m...
And they didn't even pay for the data. It was pirated. Beyond uncool in every way possible. I hope this lawsuit punches the living fuck out of his wallet. — techcrunch.com/2025/01/09/m...
Holy shit this filing update is wild! It states that: — 1. Zuckerberg gave the ok to train off copyrighted works — 2. Engineers wrote scripts to remove copyright info — 3. Meta downloaded illegal torrents to train models of — GenAi is a house of cards, I swear. — techc…
“Counsel for plaintiffs in a copyright lawsuit filed against Meta allege that Meta CEO Mark Zuckerberg gave the green light to the team behind the company's Llama AI models to use a data set of pirated ebooks and articles for training.” — https://techcrunch.com/... #news #Tech…
JUDGE THREATENS @META W/ SANCTIONS IN KADREY CASE! —NDCA Judge V. Chhabria is fed up with what he clearly considers abusive delay tactics by @Meta in seeking to seal discovery docs, calling their requests “preposterous” and obliquely, but unmistakably, warning it to stop now... […
The Order, while preserving @Meta rights to file a renewed but more limited request to seal material, also pointedly said that any overreach in doing so would result in all disputed docs being summarily unsealed by the Court! [image]
The warning comes in restrained language at the very end of a brief but blistering Order issued yesterday denying @Meta's efforts to seal docs ahead of a hearing today on heavily briefed motions and counter-filings re plaintiffs' request to file a 3d Amended Complaint: [image]
Meta Secretly Trained Its AI on a Notorious Piracy Database, Newly Unredacted Court Docs Reveal | One of the most important AI copyright legal battles just took a major turn