Court docs: internal chats suggest Meta used data from piracy site LibGen to train its Llama AI models and worked to conceal it, as Meta raced to beat rivals
A major copyright lawsuit against Meta has revealed a trove of internal communications about the company's plans to develop …
That record fed into Kadrey v. Meta, which was later permitted to proceed, before a judge ultimately issued a fair-use ruling on Meta’s book-training use that turned heavily on the plaintiffs’ arguments. The distinction between alleged acquisition practices and copyright liability remains consequential.
First-order effects
Meta faces a more detailed evidentiary record in the authors’ copyright case, while the reported chats put its internal data-governance and disclosure practices under sharper scrutiny.
Authors and rightsholders gain material that could support discovery and arguments about how Llama’s training corpus was assembled, though the documents alone do not establish infringement.
Second-order effects
Other AI developers have a clearer incentive to document dataset sourcing, approvals, and retention decisions, since internal communications can shape litigation risk alongside the underlying training use.
Publishers and licensing intermediaries gain leverage to press for paid, auditable corpus arrangements, especially where developers cannot clearly trace training material to authorized sources.
Third-order effects
If courts and counterparties continue to probe provenance separately from fair use, AI training may shift toward governed corpora with stronger audit trails rather than treating broadly available data as operationally interchangeable.
The eventual boundary remains unsettled: a favorable fair-use result in one case need not eliminate exposure tied to how material was obtained, recorded, or represented.
The trend: Generative-AI competition is turning training-data provenance into a core legal, commercial, and governance differentiator.
ah yes, Library Genesis, the piracy website ruined by TikTokers showing it off as a “viral cool life hack” to get their college textbooks for free. it was taken down shortly after. [embedded post]
I suspect a lot of what's going on with Meta is they don't have compelling products. Quest VR is very neat, but it's not setting the world on fire. Facebook's audience is ancient. Instagram is/was losing to TikTok. Threads is doing okay, but there's not much money in it. And…
Just as we all thought. Theft from top to bottom. And later in the thread there's memos where Zuckerberg's directing folks to just REMOVE copyright and ISBNs. [embedded post]
Little known fact, but putting “attorney client privilege” at the top of an email does not automatically make your email privileged. Also, as a general rule, if you cc your attorney on an email to someone who is not your attorney, the email will not then become privileged.
This retweet from December rings differently having now seen these exhibits. A lot differently. What are the chances OpenAI isn't worse? Much worse? [image]
Unsealed court documents today suggest that Meta approved the use of pirated work to train AI because they wanted to compete with other AI companies they suspected of doing the same thing. It is infuriating that governments are considering legalizing IP theft to benefit AI
Side note, this is similar to what Facebook appeared to do in 2018 after its Cambridge Analytica scandal broke wide open. It was in unsealed docs as “Project Lighthouse,” a board authorized injection of global personal datasets to train its advertising algorithms. Patterns.
wow. Upon Court order, incriminating exhibits were unsealed at 3:30am in an AI lawsuit against Meta. Once past a ‘fake privilege,’ it appears Zuckerberg approved the use of a highly controversial, pirated dataset. Note OpenAI, too? AI companies with no ethics or guardrails. /1 [i…