Hachette, Elsevier, Cengage Learning, and author Scott Turow sue Google for allegedly using millions of copyrighted books and articles to build AI models
Publishers claim the tech giant used millions of copyrighted works without permission to build its Gemini artificial intelligence modelsSee also Mediagazer
Context & Ripple Effects
The same publisher group and Scott Turow previously brought a class-action copyright case against Meta, while authors and news publishers have also challenged alleged unlicensed AI-training use by Microsoft, OpenAI, and others. This filing extends that publisher-led campaign to Google and Gemini.
The dispute follows a longer publishing-industry effort to police large-scale digital copying, including the publishers’ earlier case against Internet Archive. It puts the treatment of books and scholarly articles in AI training at the center of another major-platform lawsuit.
First-order effects
- Google must defend Gemini’s training-data practices against allegations that it used millions of copyrighted books and articles without permission.
- Hachette, Elsevier, Cengage Learning, and Turow gain another venue to press for accountability over alleged use of their works in AI development.
Second-order effects
- The parallel cases against Meta and Google increase pressure on AI developers to document training-data provenance and negotiate with book and academic publishers rather than treat their catalogs as a largely undifferentiated web-data source.
- Publishers’ licensing leverage rises if multiple major model developers face similar claims, particularly for high-value educational, scholarly, and literary material.
Third-order effects
- If courts or settlements consistently favor rightsholders, AI-model development could shift toward more formal content-licensing arrangements and clearer distinctions between authorized and allegedly pirated training corpora.
- The growing concentration of claims around a small set of major AI platforms may turn copyright exposure into a lasting competitive and cost factor for foundation-model development, though outcomes will depend on how courts assess training uses.
The trend: This is part of the widening effort by publishers and authors to establish copyright and licensing rules for the data used to train generative AI systems.