Filing: in its ChatGPT lawsuit, the NYT asked OpenAI to produce 120M user chats to analyze how often ChatGPT regurgitates its articles, after OpenAI offered 20M
Ashley Belanger / Ars Technica : Bluesky: @quinnypig.com . Forums: Slashdot and Ars OpenForum Bluesky: Corey Quinn / @quinnypig.com : I am no statistician, but can't you get meaningful signal from orders of magnitude less data? [embedded post] Forums: BeauHD / Slashdot : OpenAI Offers 20 Million User Chats In ChatGPT Lawsuit. NYT Wants 120 Million. Ars OpenForum : OpenAI offers 20 million user chats in ChatGPT lawsuit. NYT wants 120 million.
Context & Ripple Effects
The discovery dispute sits alongside OpenAI's effort to resist a preservation order for ChatGPT logs, arguing that retaining even deleted chats conflicts with user-privacy commitments. Its proposed 20 million-chat production would test alleged output reproduction at a scale far beyond conventional examples.
Later coverage records a judge requiring production of 20 million anonymized ChatGPT chat logs, making the gap between the parties' requested scopes central to how the copyright claims can be tested.
First-order effects
- The Times and OpenAI must litigate the appropriate discovery sample: the Times seeks 120 million chats to measure alleged article regurgitation, while OpenAI has offered 20 million.
- ChatGPT users' conversations become a more immediate privacy and data-governance issue because the dispute concerns production of anonymized user logs; OpenAI had already challenged the broad log-preservation requirement.
Second-order effects
- The eventual sample and anonymization standards can determine what evidence each side can use to characterize output behavior, shaping leverage in this copyright case.
- Other publishers and AI providers will watch whether large-scale user-output records become a practical route to proving or contesting claims that models reproduce protected reporting.
Third-order effects
- If courts increasingly treat user prompts and outputs as necessary evidence in generative-AI copyright cases, privacy-preserving discovery processes may become a recurring operating requirement for consumer AI services.
- The dispute reinforces a broader shift from abstract training-data arguments toward evidence about model behavior in use, which could influence future content-licensing and permission-boundary negotiations.
The trend: Generative-AI copyright litigation is increasingly testing whether platform-scale user-output data can establish the line between useful summarization and protected-content reproduction.