Following similar lawsuits, Sarah Silverman and two other authors sue OpenAI and Meta, claiming LLaMA and ChatGPT were trained on copyright-infringing material
Comedian and author Sarah Silverman, as well as authors Christopher Golden and Richard Kadrey — are suing OpenAI and Meta each …
Context & Ripple Effects
This complaint joined the early wave of author challenges to generative-AI training, bringing the same core allegation to both OpenAI's ChatGPT and Meta's Llama. It matters because it put the provenance of model-training material—not just model outputs—at the center of the dispute.
The case became a durable test of that theory: most claims against Meta were later dismissed, yet the remaining dispute advanced toward a federal court allowing the Kadrey v. Meta case to proceed and a later trial focused on alleged LibGen use. OpenAI also sought dismissal of similar book-author copyright claims, underscoring the parallel legal tracks facing major model developers.
First-order effects
- Sarah Silverman, Christopher Golden, and Richard Kadrey put OpenAI and Meta on notice that their alleged use of copyright-infringing training material could create direct copyright exposure for ChatGPT and Llama.
- The suits force both companies to defend how training datasets were assembled and whether that use is legally protected, rather than treating the dispute solely as one over generated outputs.
Second-order effects
- Authors' representatives gain a vehicle to seek evidence about training-data provenance; that pressure later produced authors' access to review OpenAI training data in related litigation.
- Other AI developers and dataset suppliers face stronger incentives to document sourcing and prepare fair-use defenses as book-rights holders test similar claims.
Third-order effects
- The litigation could help determine whether large-scale ingestion of copyrighted works is treated primarily as protected model development or as a licensable input—a boundary with consequences for training-data markets.
- If courts demand more traceability or rights clearance, model builders may shift toward more auditable datasets and negotiated content access; the eventual legal standard remains unsettled.
The trend: Generative-AI copyright disputes are turning training-data provenance into a core constraint on how foundation models are built and commercialized.