/
Navigation
Chronicles
Browse all articles
Explore
Semantic exploration
Research
Entity momentum
Nexus
Correlations & relationships
Story Arc
Topic evolution
Drift Map
Semantic trajectory animation
Posts
Analysis & commentary
Pulse API
Tech news intelligence API
Browse
Entities
Companies, people, products, technologies
Domains
Browse by publication source
Handles
Browse by social media handle
Detection
Concept Search
Semantic similarity search
High Impact Stories
Top coverage by position
Sentiment Analysis
Positive/negative coverage
Anomaly Detection
Unusual coverage patterns
Analysis
Rivalry Report
Compare two entities head-to-head
Semantic Pivots
Narrative discontinuities
Crisis Response
Event recovery patterns
Connected
Search: /
Command: ⌘K
Embeddings: large
TEXXR

Chronicles

The story behind the story

days · browse · Enter similar · o open

In a paper, media mogul Tim O'Reilly and economist Ilan Strauss say OpenAI likely trained GPT-4o on paywalled O'Reilly Media books without a licensing agreement

OpenAI has been accused by many parties of training its AI on copyrighted content sans permission.

TechCrunch Kyle Wiggers

Context & Ripple Effects

The allegation extends a running dispute over whether OpenAI secured permission for material used in training. Earlier coverage recorded Dow Jones saying it had no training agreement for Wall Street Journal reporting, while reporting also described OpenAI’s use of transcribed YouTube material for GPT-4 training.

It matters because O’Reilly and Strauss frame the issue around a paywalled professional publishing catalog and argue that model providers can build arrangements that compensate rights holders, rather than treating access to web-derived text as the default.

First-order effects

  • The paper puts O’Reilly Media’s books into the training-data provenance debate and increases pressure on OpenAI to address a specific claim concerning GPT-4o’s inputs.
  • For publishers, it adds another concrete example alongside Dow Jones’s earlier claim that no WSJ training deal existed, strengthening the case for clarity on whether licenses were obtained.

Second-order effects

  • Publishers with paid archives may have more incentive to seek licensing terms or disclosures from model providers, particularly where their catalogs are suited to technical or professional AI use.
  • OpenAI and rivals face a sharper commercial trade-off: negotiate access to high-value corpora or defend training practices amid recurring questions about data sourcing, including reported transcription of YouTube material for GPT-4 training.

Third-order effects

  • If claims tied to identifiable paid catalogs keep accumulating, training-data access could shift from an opaque collection problem toward a governed-content procurement market.
  • That shift would make documentation of corpus rights and creator compensation a more central competitive and governance issue, though the article does not establish how OpenAI will respond or whether a license was required.

The trend: Generative AI is moving toward a contest over whether valuable content should be treated as freely usable training input or as licensed infrastructure for model development.