In a paper, media mogul Tim O'Reilly and economist Ilan Strauss say OpenAI likely trained GPT-4o on paywalled O'Reilly Media books without a licensing agreement
OpenAI has been accused by many parties of training its AI on copyrighted content sans permission.
Context & Ripple Effects
The allegation extends a running dispute over whether OpenAI secured permission for material used in training. Earlier coverage recorded Dow Jones saying it had no training agreement for Wall Street Journal reporting, while reporting also described OpenAI’s use of transcribed YouTube material for GPT-4 training.
It matters because O’Reilly and Strauss frame the issue around a paywalled professional publishing catalog and argue that model providers can build arrangements that compensate rights holders, rather than treating access to web-derived text as the default.
First-order effects
- The paper puts O’Reilly Media’s books into the training-data provenance debate and increases pressure on OpenAI to address a specific claim concerning GPT-4o’s inputs.
- For publishers, it adds another concrete example alongside Dow Jones’s earlier claim that no WSJ training deal existed, strengthening the case for clarity on whether licenses were obtained.
Second-order effects
- Publishers with paid archives may have more incentive to seek licensing terms or disclosures from model providers, particularly where their catalogs are suited to technical or professional AI use.
- OpenAI and rivals face a sharper commercial trade-off: negotiate access to high-value corpora or defend training practices amid recurring questions about data sourcing, including reported transcription of YouTube material for GPT-4 training.
Third-order effects
- If claims tied to identifiable paid catalogs keep accumulating, training-data access could shift from an opaque collection problem toward a governed-content procurement market.
- That shift would make documentation of corpus rights and creator compensation a more central competitive and governance issue, though the article does not establish how OpenAI will respond or whether a license was required.
The trend: Generative AI is moving toward a contest over whether valuable content should be treated as freely usable training input or as licensed infrastructure for model development.