/
Navigation
Chronicles
Browse all articles
Explore
Semantic exploration
Research
Entity momentum
Nexus
Correlations & relationships
Story Arc
Topic evolution
Drift Map
Semantic trajectory animation
Posts
Analysis & commentary
Pulse API
Tech news intelligence API
Browse
Entities
Companies, people, products, technologies
Domains
Browse by publication source
Handles
Browse by social media handle
Detection
Concept Search
Semantic similarity search
High Impact Stories
Top coverage by position
Sentiment Analysis
Positive/negative coverage
Anomaly Detection
Unusual coverage patterns
Analysis
Rivalry Report
Compare two entities head-to-head
Semantic Pivots
Narrative discontinuities
Crisis Response
Event recovery patterns
Connected
Search: /
Command: ⌘K
Embeddings: large
TEXXR

Chronicles

The story behind the story

days · browse · Enter similar · o open

Filing: in its ChatGPT lawsuit, the NYT asked OpenAI to produce 120M user chats to analyze how often ChatGPT regurgitates its articles, after OpenAI offered 20M

Ashley Belanger / Ars Technica : Bluesky: @quinnypig.com . Forums: Slashdot and Ars OpenForum Bluesky: Corey Quinn / @quinnypig.com : I am no statistician, but can't you get meaningful signal from orders of magnitude less data? [embedded post] Forums: BeauHD / Slashdot : OpenAI Offers 20 Million User Chats In ChatGPT Lawsuit. NYT Wants 120 Million. Ars OpenForum : OpenAI offers 20 million user chats in ChatGPT lawsuit. NYT wants 120 million.

Ars Technica Ashley Belanger

Context & Ripple Effects

The discovery dispute sits alongside OpenAI's effort to resist a preservation order for ChatGPT logs, arguing that retaining even deleted chats conflicts with user-privacy commitments. Its proposed 20 million-chat production would test alleged output reproduction at a scale far beyond conventional examples.

Later coverage records a judge requiring production of 20 million anonymized ChatGPT chat logs, making the gap between the parties' requested scopes central to how the copyright claims can be tested.

First-order effects

  • The Times and OpenAI must litigate the appropriate discovery sample: the Times seeks 120 million chats to measure alleged article regurgitation, while OpenAI has offered 20 million.
  • ChatGPT users' conversations become a more immediate privacy and data-governance issue because the dispute concerns production of anonymized user logs; OpenAI had already challenged the broad log-preservation requirement.

Second-order effects

  • The eventual sample and anonymization standards can determine what evidence each side can use to characterize output behavior, shaping leverage in this copyright case.
  • Other publishers and AI providers will watch whether large-scale user-output records become a practical route to proving or contesting claims that models reproduce protected reporting.

Third-order effects

  • If courts increasingly treat user prompts and outputs as necessary evidence in generative-AI copyright cases, privacy-preserving discovery processes may become a recurring operating requirement for consumer AI services.
  • The dispute reinforces a broader shift from abstract training-data arguments toward evidence about model behavior in use, which could influence future content-licensing and permission-boundary negotiations.

The trend: Generative-AI copyright litigation is increasingly testing whether platform-scale user-output data can establish the line between useful summarization and protected-content reproduction.

Discussion

  • @quinnypig.com Corey Quinn on bluesky
    I am no statistician, but can't you get meaningful signal from orders of magnitude less data? [embedded post]