/
Navigation
Chronicles
Browse all articles
Explore
Semantic exploration
Research
Entity momentum
Nexus
Correlations & relationships
Story Arc
Topic evolution
Drift Map
Semantic trajectory animation
Posts
Analysis & commentary
Pulse API
Tech news intelligence API
Browse
Entities
Companies, people, products, technologies
Domains
Browse by publication source
Handles
Browse by social media handle
Detection
Concept Search
Semantic similarity search
High Impact Stories
Top coverage by position
Sentiment Analysis
Positive/negative coverage
Anomaly Detection
Unusual coverage patterns
Analysis
Rivalry Report
Compare two entities head-to-head
Semantic Pivots
Narrative discontinuities
Crisis Response
Event recovery patterns
Connected
Search: /
Command: ⌘K
Embeddings: large
TEXXR

Chronicles

The story behind the story

days · browse · Enter similar · o open

Court docs: internal chats suggest Meta used data from piracy site LibGen to train its Llama AI models and worked to conceal it, as Meta raced to beat rivals

A major copyright lawsuit against Meta has revealed a trove of internal communications about the company's plans to develop …

The Verge

Context & Ripple Effects

The disclosures add alleged concealment and competitive urgency to earlier reporting that Zuckerberg approved a Llama training effort using LibGen-sourced material. They make the provenance of model-training data—not only the model’s output—a central issue in the dispute.

That record fed into Kadrey v. Meta, which was later permitted to proceed, before a judge ultimately issued a fair-use ruling on Meta’s book-training use that turned heavily on the plaintiffs’ arguments. The distinction between alleged acquisition practices and copyright liability remains consequential.

First-order effects

  • Meta faces a more detailed evidentiary record in the authors’ copyright case, while the reported chats put its internal data-governance and disclosure practices under sharper scrutiny.
  • Authors and rightsholders gain material that could support discovery and arguments about how Llama’s training corpus was assembled, though the documents alone do not establish infringement.

Second-order effects

  • Other AI developers have a clearer incentive to document dataset sourcing, approvals, and retention decisions, since internal communications can shape litigation risk alongside the underlying training use.
  • Publishers and licensing intermediaries gain leverage to press for paid, auditable corpus arrangements, especially where developers cannot clearly trace training material to authorized sources.

Third-order effects

  • If courts and counterparties continue to probe provenance separately from fair use, AI training may shift toward governed corpora with stronger audit trails rather than treating broadly available data as operationally interchangeable.
  • The eventual boundary remains unsettled: a favorable fair-use result in one case need not eliminate exposure tied to how material was obtained, recorded, or represented.

The trend: Generative-AI competition is turning training-data provenance into a core legal, commercial, and governance differentiator.

Discussion

  • @qwertybro @qwertybro on bluesky
    ah yes, Library Genesis, the piracy website ruined by TikTokers showing it off as a “viral cool life hack” to get their college textbooks for free. it was taken down shortly after.  [embedded post]
  • @aaronrosspowell.com Aaron Ross Powell on bluesky
    I suspect a lot of what's going on with Meta is they don't have compelling products.  Quest VR is very neat, but it's not setting the world on fire.  Facebook's audience is ancient.  Instagram is/was losing to TikTok.  Threads is doing okay, but there's not much money in it.  And…
  • @oheysteenz @oheysteenz on bluesky
    Just as we all thought.  Theft from top to bottom.  And later in the thread there's memos where Zuckerberg's directing folks to just REMOVE copyright and ISBNs.  [embedded post]
  • @daniel_271828 Daniel Eth on x
    Little known fact, but putting “attorney client privilege” at the top of an email does not automatically make your email privileged. Also, as a general rule, if you cc your attorney on an email to someone who is not your attorney, the email will not then become privileged.
  • @jason_kint Jason Kint on x
    This retweet from December rings differently having now seen these exhibits. A lot differently. What are the chances OpenAI isn't worse? Much worse? [image]
  • @ednewtonrex Ed Newton-Rex on x
    Unsealed court documents today suggest that Meta approved the use of pirated work to train AI because they wanted to compete with other AI companies they suspected of doing the same thing. It is infuriating that governments are considering legalizing IP theft to benefit AI
  • @jason_kint Jason Kint on x
    Side note, this is similar to what Facebook appeared to do in 2018 after its Cambridge Analytica scandal broke wide open. It was in unsealed docs as “Project Lighthouse,” a board authorized injection of global personal datasets to train its advertising algorithms. Patterns.
  • @jason_kint Jason Kint on x
    wow. Upon Court order, incriminating exhibits were unsealed at 3:30am in an AI lawsuit against Meta. Once past a ‘fake privilege,’ it appears Zuckerberg approved the use of a highly controversial, pirated dataset. Note OpenAI, too? AI companies with no ethics or guardrails. /1 [i…