/
Navigation
Chronicles
Browse all articles
Explore
Semantic exploration
Research
Entity momentum
Nexus
Correlations & relationships
Story Arc
Topic evolution
Drift Map
Semantic trajectory animation
Posts
Analysis & commentary
Pulse API
Tech news intelligence API
Browse
Entities
Companies, people, products, technologies
Domains
Browse by publication source
Handles
Browse by social media handle
Detection
Concept Search
Semantic similarity search
High Impact Stories
Top coverage by position
Sentiment Analysis
Positive/negative coverage
Anomaly Detection
Unusual coverage patterns
Analysis
Rivalry Report
Compare two entities head-to-head
Semantic Pivots
Narrative discontinuities
Crisis Response
Event recovery patterns
Connected
Search: /
Command: ⌘K
Embeddings: large
TEXXR

Chronicles

The story behind the story

days · browse · Enter similar · o open

Dow Jones' general counsel says OpenAI lacks a deal to use WSJ reporting to train its AI; a source says CNN plans to ask OpenAI to pay to license its content

Major news outlets have begun criticizing OpenAI and its ChatGPT software, saying the lab is using their articles to train …

Bloomberg Gerry Smith

Context & Ripple Effects

The dispute marks an early split between publishers seeking payment for AI use and OpenAI's access to their reporting. That split later became explicit in the New York Times copyright suit and in licensing talks with major publishers.

CNN's stated intent to seek payment foreshadowed the later report that it was among outlets discussing text, video, and image licensing with OpenAI. The key issue is not merely chatbot output, but whether newsroom archives become a paid model-development input.

First-order effects

  • Dow Jones can withhold consent for Wall Street Journal reporting and publicly establish that OpenAI has no training-use agreement with it.
  • CNN gains a defined licensing position to put to OpenAI: payment for use of its content rather than uncompensated access.

Second-order effects

  • OpenAI's negotiations with publishers become more consequential, as the absence of agreements with prominent outlets gives publishers leverage to press for standardized commercial terms.
  • Other newspaper groups have a clearer litigation alternative to licensing, a path later taken by the Alden-owned daily newspapers against OpenAI and Microsoft.

Third-order effects

  • News archives are being recast from material platforms can ingest by default into a licensable input for model training, with publisher consent becoming a competitive procurement issue.
  • The parallel emergence of publisher deals and copyright claims points toward a bifurcated market: negotiated access for some outlets and court-defined boundaries for others.

The trend: Generative-AI developers and news publishers are moving toward a market in which high-value reporting is priced as training and product input rather than treated as freely available web material.

Discussion

  • @business @business on x
    Major news outlets have begun criticizing OpenAI and its ChatGPT software, saying the lab is using their articles to train its artificial intelligence tool without paying them https://www.bloomberg.com/...
  • @tolles Chris Tolles on x
    @fpmarconi Their Robots.txt spell out the crawling policies. You don't need an agreement to crawl a site. Seems like a speculative accusation without understanding of the way the web works unless you have something more solid
  • @simonowens Simon Owens on x
    I'm highly skeptical that you can force a company to pay you just to train its AI on your freely-available content. If that's the case, then companies like Google would be forced to pay for scanning the entire internet every few hour just to update its search algorithms. https://…
  • @gerryfsmith Gerry Smith on x
    WSJ and CNN are concerned that OpenAI is using their stories to train its artificial intelligence technology without paying them. “We take the misuse of our journalists' work seriously, and are reviewing this situation.” https://www.bloomberg.com/...
  • @fpmarconi Francesco Marconi on x
    Here's the prompt I used: “Which specific news sources was chatGPT trained on? Provide a list of the top news sources in your database.”
  • @sub8u @sub8u on x
    Then: Google is making money by crawling through our website. Now: Large language models are making money by using our articles to train the AI systems Forever: Everybody is making money off us. We deserve to be be paid. https://twitter.com/...
  • @tolles Chris Tolles on x
    @Techmeme @gerryfsmith Guessing this is going to go to court and the “right to read” for machines litigated. Right now crawling public info is legally solid (and if that violates a TOS, so far courts have said that's ok, from what I can tell)
  • @carnage4life Dare Obasanjo on x
    If you think the media were mad at Facebook and Google News for replacing them as a source of news, this will be nothing compared to how much they're going to go after OpenAI for “stealing” their content. This is Stable Diffusion versus artists all over. https://www.bloomberg.com…
  • @rafat @rafat on x
    @tolles ... My only ask would be to add an option to “untrain” their LLM on a publisher's corpus of content if that is what the publisher wants. Seems a reasonable request. Robots.txt does the same thing for search engines.
  • @fpmarconi Francesco Marconi on x
    ChatGPT is not able to generate original reporting but it can be helpful for repurposing existing content — for example localizing news stories. Human validation is still required to check for potential mistakes and data hallucinations. https://www.ft.com/...
  • @tolles Chris Tolles on x
    @fpmarconi Ok. I will concede this is a greyer area from a pure tos standpoint (and withdraw my accusation of libel) but when there is a discrepancy between posted TOS and robots.txt, in practice, a crawler is engaging in a protected action. https://law.stackexchange.com/ ...
  • @sarhan_ Sarhan on x
    Articles are written to be read, either by humans, search engines, or AI.. What's wrong with that? https://twitter.com/...
  • @martijnrasser Martijn Rasser on x
    Part of a trend. AP started using AI in 2016 to write recaps of minor league baseball games. From this article: “Thomson Reuters has used an in-house program since 2018 to sift through information such as market data to find patterns for reporters.” https://www.ft.com/...
  • @davidgaw David Gaw on x
    @Techmeme @gerryfsmith A license isn't required to read information on the Internet, whether the reader is a person or an AI. If you want to license your content, put it behind a paywall.
  • @fpmarconi Francesco Marconi on x
    News organizations have a new audience consuming more content than anyone else: machines. It seems fair for publishers to be compensated if their content is used to train someone else's AI. https://www.bloomberg.com/...
  • @groomb Brian Groom on x
    Reach, Mirror, Express regional publisher, explores whether ChatGPT could help journalists write short news stories. Working group to examine how tool might be used to assist human reporters compiling coverage of topics such as local weather and traffic. https://www.ft.com/...
  • @fpmarconi Francesco Marconi on x
    @tolles The legal debate between AI generators and content rights-holders will depend on the interpretation of US Fair Use doctrine and “transformative use” — and whether systems like GPT store this data in their databases.
  • @marcusschuler Marcus Schuler on x
    The first media companies are demanding license fees from OpenAI for the use of their content. And they are right to do so. Because with #ChatGPT and Co., the content is extracted from the pages of the content providers, i.e., media companies. https://www.bloomberg.com/...
  • @newley Newley Purnell on x
    “Anyone who wants to use the work of Wall Street Journal journalists to train artificial intelligence should be properly licensing the rights to do so from Dow Jones...Dow Jones does not have such a deal with OpenAI.” https://www.bloomberg.com/...