A US judge says OpenAI must face a claim of violating CA unfair competition law by using copyrighted books, but dismisses some claims like DMCA violations
Context & Ripple Effects
This early ruling separates a state-law challenge over training material from the dismissed DMCA theory, leaving OpenAI with a narrower but still consequential copyright-related dispute. Later coverage shows that the procedural picture remains uneven: the New York Times' core infringement claims survived OpenAI's dismissal bid, while another publisher suit was dismissed for insufficiently shown harm.
The wider record has also begun to distinguish between uses of books rather than treating AI training as one legal act. A later Anthropic ruling found training on copyrighted books fair use but not the retention of pirated copies, underscoring why the source and handling of training corpora matter.
First-order effects
- OpenAI must continue defending the California unfair-competition claim tied to its use of copyrighted books, preserving discovery and litigation exposure on that theory.
- Dismissal of the DMCA allegations narrows the case and removes one asserted route to liability at this stage.
Second-order effects
- Copyright plaintiffs gain a viable state-law pleading path alongside direct infringement claims, while AI developers have added incentive to document how training material was acquired and retained.
- The split treatment of claims makes case-by-case outcomes more likely: plaintiffs will need to connect alleged conduct to a cognizable harm, as the Raw Story and AlterNet case showed when harm allegations fell short.
Third-order effects
- If courts continue to differentiate lawful training from unlawful acquisition or storage, model developers' data-governance practices could become as important as the models' outputs in litigation risk.
- The durable question is shifting toward the permission boundary for commercial AI datasets, though the available rulings do not yet establish a uniform rule across claims or jurisdictions.
The trend: Generative-AI copyright litigation is developing into a granular test of dataset provenance, retained copies, and the fit between legacy legal theories and model training.