Dow Jones' general counsel says OpenAI lacks a deal to use WSJ reporting to train its AI; a source says CNN plans to ask OpenAI to pay to license its content
Major news outlets have begun criticizing OpenAI and its ChatGPT software, saying the lab is using their articles to train …
CNN's stated intent to seek payment foreshadowed the later report that it was among outlets discussing text, video, and image licensing with OpenAI. The key issue is not merely chatbot output, but whether newsroom archives become a paid model-development input.
First-order effects
Dow Jones can withhold consent for Wall Street Journal reporting and publicly establish that OpenAI has no training-use agreement with it.
CNN gains a defined licensing position to put to OpenAI: payment for use of its content rather than uncompensated access.
Second-order effects
OpenAI's negotiations with publishers become more consequential, as the absence of agreements with prominent outlets gives publishers leverage to press for standardized commercial terms.
Other newspaper groups have a clearer litigation alternative to licensing, a path later taken by the Alden-owned daily newspapers against OpenAI and Microsoft.
Third-order effects
News archives are being recast from material platforms can ingest by default into a licensable input for model training, with publisher consent becoming a competitive procurement issue.
The parallel emergence of publisher deals and copyright claims points toward a bifurcated market: negotiated access for some outlets and court-defined boundaries for others.
The trend: Generative-AI developers and news publishers are moving toward a market in which high-value reporting is priced as training and product input rather than treated as freely available web material.
Major news outlets have begun criticizing OpenAI and its ChatGPT software, saying the lab is using their articles to train its artificial intelligence tool without paying them https://www.bloomberg.com/...
@fpmarconi Their Robots.txt spell out the crawling policies. You don't need an agreement to crawl a site. Seems like a speculative accusation without understanding of the way the web works unless you have something more solid
I'm highly skeptical that you can force a company to pay you just to train its AI on your freely-available content. If that's the case, then companies like Google would be forced to pay for scanning the entire internet every few hour just to update its search algorithms. https://…
WSJ and CNN are concerned that OpenAI is using their stories to train its artificial intelligence technology without paying them. “We take the misuse of our journalists' work seriously, and are reviewing this situation.” https://www.bloomberg.com/...
Then: Google is making money by crawling through our website. Now: Large language models are making money by using our articles to train the AI systems Forever: Everybody is making money off us. We deserve to be be paid. https://twitter.com/...
@Techmeme @gerryfsmith Guessing this is going to go to court and the “right to read” for machines litigated. Right now crawling public info is legally solid (and if that violates a TOS, so far courts have said that's ok, from what I can tell)
If you think the media were mad at Facebook and Google News for replacing them as a source of news, this will be nothing compared to how much they're going to go after OpenAI for “stealing” their content. This is Stable Diffusion versus artists all over. https://www.bloomberg.com…
@tolles ... My only ask would be to add an option to “untrain” their LLM on a publisher's corpus of content if that is what the publisher wants. Seems a reasonable request. Robots.txt does the same thing for search engines.
ChatGPT is not able to generate original reporting but it can be helpful for repurposing existing content — for example localizing news stories. Human validation is still required to check for potential mistakes and data hallucinations. https://www.ft.com/...
@fpmarconi Ok. I will concede this is a greyer area from a pure tos standpoint (and withdraw my accusation of libel) but when there is a discrepancy between posted TOS and robots.txt, in practice, a crawler is engaging in a protected action. https://law.stackexchange.com/ ...
Part of a trend. AP started using AI in 2016 to write recaps of minor league baseball games. From this article: “Thomson Reuters has used an in-house program since 2018 to sift through information such as market data to find patterns for reporters.” https://www.ft.com/...
@Techmeme @gerryfsmith A license isn't required to read information on the Internet, whether the reader is a person or an AI. If you want to license your content, put it behind a paywall.
News organizations have a new audience consuming more content than anyone else: machines. It seems fair for publishers to be compensated if their content is used to train someone else's AI. https://www.bloomberg.com/...
Reach, Mirror, Express regional publisher, explores whether ChatGPT could help journalists write short news stories. Working group to examine how tool might be used to assist human reporters compiling coverage of topics such as local weather and traffic. https://www.ft.com/...
@tolles The legal debate between AI generators and content rights-holders will depend on the interpretation of US Fair Use doctrine and “transformative use” — and whether systems like GPT store this data in their databases.
The first media companies are demanding license fees from OpenAI for the use of their content. And they are right to do so. Because with #ChatGPT and Co., the content is extracted from the pages of the content providers, i.e., media companies. https://www.bloomberg.com/...
“Anyone who wants to use the work of Wall Street Journal journalists to train artificial intelligence should be properly licensing the rights to do so from Dow Jones...Dow Jones does not have such a deal with OpenAI.” https://www.bloomberg.com/...