Reddit plans to charge companies to access its API, which many have used to train AI tools; the API will stay free for developers building apps for Reddit users
The internet site has long been a forum for discussion on a huge variety of topics, and companies like Google and OpenAI have been using it in their A.I. projects.
New York TimesMike Isaac
Context & Ripple Effects
Reddit’s distinction between commercial AI use and user-facing app development turned its discussion archive into a product rather than a freely available input. Follow-up coverage tied the policy to LLMs increasing the value of Reddit data and the company’s IPO preparations, while Reddit later reached a reported $60M-a-year data-access agreement with Google.
The move matters because it establishes an early boundary around who may reuse community-generated content at scale: developers serving Reddit users retain access, while model builders and other commercial users face a licensing path.
First-order effects
AI companies and other commercial users of Reddit’s API must negotiate paid access instead of treating Reddit content as a no-cost training resource.
Developers building apps for Reddit users remain exempt, preserving the ecosystem of user-facing integrations while separating it from commercial data extraction.
Second-order effects
Paid access gives Reddit a mechanism to convert its data into recurring commercial revenue; the reported Google agreement shows how that mechanism could become a direct licensing channel.
Model developers reliant on large public-discussion datasets face higher acquisition costs and incentives to secure comparable data partnerships rather than depend on unrestricted APIs.
Third-order effects
If replicated by other content platforms, publicly accessible community data could shift from an informal AI-training commons toward negotiated, platform-controlled inputs.
That transition may create a persistent trade-off for publishers and platforms: licensing AI access can monetize content, but later coverage that outlets were considering limiting Google’s AI access as referral traffic fell suggests distribution effects can complicate those deals.
The trend: This is an early example of AI training data being commercialized through controlled API access rather than treated as an open web byproduct.
some news: Reddit will begin charging the biggest companies for API access, which has been used historically to train the coming wave of LLMs and artificially intelligent programs “It's a good time for us to tighten things up,” CEO Steve Huffman said. https://www.nytimes.com/...
Sigh... as tho reddit cannot be scraped without API. This move to close the internet down to developers who want to enrich ecosystem (Twitter, now Reddit) is a very bad outcome, and I wonder how much of a solution it really offers. https://twitter.com/...
“The Reddit corpus of data is really valuable,” Reddit's founder and CEO, Steve Huffman, told @MikeIsaac. “But we don't need to give all of that value to some of the largest companies in the world for free.” https://www.nytimes.com/...
I, for one, think it's a good idea to build a titanomarchy of new world-striding spider gods from the collective minds of the lamest, most annoying people to ever live. https://twitter.com/...
Gonna be fascinating. AI folks say they shouldn't get charged to look at things on the web - but then some of them do pay to ingest stuff: e.g. OpenAI pays shutterstock. Meanwhile many content makers/owners will be dismayed to find how little their content is worth. https://twitt…
Gonna be interesting 5 years from now when we see a bunch of LLMs (or whatever we then call it) whose training data basically stops at midway-through-2023 Reddit. https://twitter.com/...
“More than any other place on the internet, Reddit is a home for authentic conversation. There's a lot of stuff on the site that you'd only ever say in therapy, or A.A., or never at all.” https://twitter.com/...