In 2026, Reddit said its content-licensing deals with Google and OpenAI accounted for roughly 10% of revenue. Its core search reached more than 70M weekly unique users. The web had spent two decades treating that archive as exhaust before Reddit attached a visible price to it.

Key takeaways

  • Reddit said content-licensing deals with Google and OpenAI accounted for roughly 10% of its revenue in 2026.
  • Stack Overflow’s CEO said in 2023 that the company planned to charge major AI developers for access to 50 million questions and answers.
  • An analysis of 1,800 AI datasets found that roughly 70% lacked a stated license or used labels more permissive than their creators intended.
  • Stack Overflow’s annual revenue doubled to $115 million after ChatGPT, driven by enterprise products and AI licensing.
  • Digg’s 2026 reboot followed an invite-only period involving 67,000 users.

Reddit’s advertising revenue grew 60% year over year in Q4, so licensing remains secondary. Reddit Answers reached 6M weekly users, leaving core search more than eleven times as large as the new answer interface. That gap measures adoption, not how much the record is worth to buyers.

Google, OpenAI, and other buyers now value community platforms for archives whose owners can authenticate material, update it, and set access terms. Reddit and Stack Overflow must sell controlled access without letting AI answers destroy the referrals, human contributions, and moderation judgments that make those archives worth buying.

The archive became infrastructure before it became a product

Common Crawl’s 9.5PB archive, dating to 2008, made freely collected web-scale history a default input for model builders. It converted unruly public pages into machine-scale training material without requiring builders to negotiate separately with every site.

Community archives carry another layer. Reddit and Stack Overflow preserve discussions, replies, disagreement, votes, reputation, recency, and moderation decisions. The words are only one part of a thread; its surrounding judgments supply context.

Model trainers consume large historical samples. Retrieval systems need current, queryable slices, and answer engines combine both. When an AI system needs information beyond its training data and context window, it must reach an external source, giving a platform-owned archive an advantage over copied pages.

Stack Overflow identified the commercial opening early. Its CEO said in 2023 that the company planned to charge major AI developers for access to 50M questions and answers. Reddit later licensed content to Google and OpenAI. Those deals established organized knowledge as a negotiated input for AI systems.

Platform owners make archive scale usable when they can attest that a discussion came from their system, preserve ranking and moderation signals, update the corpus, and specify access terms. A copied dataset may retain the words while losing the judgments around them. Buyers discount inventory missing those signals.

One corpus now supports four access markets

A community platform now makes several access decisions with different buyers, time horizons, and failure modes.

Access product Primary customer Economic function Constraint
Corpus license Model developers Prices historical material and defined usage rights Provenance and consent
Metered crawler or API access Search engines and agents Prices fresh, machine-readable retrieval Enforcement and crawler evasion
First-party search and answers Users Keeps queries, context, and discovery inside the platform Answer quality and trust
Advertiser-linked retrieval Advertisers and shoppers Connects community recommendations to commercial intent Separation of recommendation from sponsorship

A contract must distinguish those rights. A training license does not necessarily permit continuous retrieval. Crawler access does not necessarily permit an answer engine to reproduce results. An answer product does not automatically carry advertising rights. The old web bundled reading, indexing, referral, and monetization into one informal bargain; AI has split that bargain into separately negotiated rights.

At the protocol layer, Cloudflare’s Pay per Crawl marketplace lets sites charge AI crawlers per visit, while new Cloudflare sites block those crawlers by default. Cloudflare charges for passage at the gate to archives it does not own.

Reddit operates farther up the stack. The company is testing AI search results that pair community recommendations with products from advertising partners. A bulk license monetizes history. Metered access monetizes freshness. First-party search captures the query. Advertising captures commercial intent. Each product draws from the same archive but needs a different contract.

AI search stopped paying the archive back

Google and other search engines once converted old discussions into new visits. By 2017, search drove 34.8% of referral traffic, compared with 25.6% from social. A useful Reddit thread or Stack Overflow answer could rank for years, attracting readers, contributors, and advertising impressions long after publication.

AI answer interfaces alter the exchange. TollBit compared referrals across 160 websites. Similarweb separately measured the rise in news searches that ended without a click after Google introduced AI Overviews.

less referral traffic from AI search engines than from Google Search across TollBit’s sample
news searches producing no click-through, May 2024 to May 2025

The samples cover different sites and measures, but both show answer interfaces keeping more utility and returning fewer visits. Google and OpenAI buy licensed access while operating interfaces that can satisfy users away from the source.

Reddit has already reported choppy search referrals. Steve Huffman also said the company would move away from r/popular, historically the default feed for new users, in favor of more personalized feeds. Reddit is moving discovery onto its own feeds as external discovery becomes less dependable.

Its own answer surface gives Reddit control over ranking, attribution, updates, and advertising. Adoption still depends on whether the interface preserves the specificity users came to Reddit to find.

A contract matters only when a platform can police it

Reddit, Yahoo, Medium, Quora, O’Reilly, wikiHow, Ziff Davis, and others adopted the Really Simple Licensing standard to express terms for AI scraping. The standard makes the public-data permission boundary machine-readable. Platform owners still have to enforce it.

Reddit sued Anthropic, alleging that the company accessed Reddit more than 100,000 times after saying it had stopped. Reddit later sued Perplexity AI and data-scraping firms over alleged copyright-control violations. Those cases expose the gap between publishing terms and enforcing them.

Sellers must also establish provenance. An analysis of 1,800 AI datasets found that roughly 70% either lacked a stated license or carried labels more permissive than creators intended. Seven dataset sellers subsequently formed the Dataset Providers Alliance around ethical sourcing because buyers cannot value a right that the seller cannot establish.

Community archives now face the same security problem as the open-AI commons becoming critical infrastructure. Provenance, permissions, and access controls become product requirements once other systems depend on the resource.

Moderators produce what buyers think they’re licensing

Reddit moderators say AI-generated low-quality posts are eroding authenticity. They also warn that models may train on synthetic material produced by earlier models. A model can generate text that enters a community, gains visibility through community signals, and returns to a later model as apparent evidence of human judgment.

A viral Reddit post from a purported developer illustrated the problem. The post alleged misconduct by a major food-delivery app, appeared AI-generated, gained broad traction, and prompted denials from Uber and DoorDash. Popularity metadata can amplify synthetic claims before provenance catches up.

Reddit has expanded testing of Rules Hub, its LLM-powered moderation tools. The company uses AI at both ends of the archive: retrieval exposes material, while moderation attempts to preserve its quality. Human moderators still provide the accountable layer. They interpret community rules, review ambiguous cases, and stand behind judgments an automated system cannot own. Their decisions determine which contributions remain visible and which signals buyers later inherit.

Digg made that relationship explicit in its 2026 reboot. After an invite-only period involving 67,000 users, Kevin Rose and Alexis Ohanian positioned open communities, transparent moderation actions, and a public algorithm as product features. Its pitch made the governance of the feed and archive visible.

Stack Overflow shows revenue can rise as questions vanish

Stack Overflow’s annual revenue doubled to $115M after ChatGPT, driven by enterprise products and AI licensing. In December 2025, its monthly question volume fell to its 2008 level.

Stack Overflow monetized its accumulated stock while the flow of new human material weakened. The figures alone cannot establish why participation fell.

Historical questions remain useful for model training and stable technical problems. Search and answer systems face a stricter test. Software libraries, APIs, and accepted practices change. A system promising current information needs new questions, corrections, votes, and moderation decisions alongside the old archive.

Stack Overflow’s enterprise tools and data rights supported revenue even as participation fell. Those products cannot manufacture the human activity that keeps a technical corpus current.

A training contract prices accumulated material. Search and answer products price current usefulness. Moderation protects the spread between copied text and trusted knowledge. Calling all three “data” is convenient in an earnings release and expensive everywhere else.

Frequently asked questions

How much did Google and OpenAI each pay Reddit for content access?

The piece reports only that the licensing deals together represented roughly 10% of Reddit revenue. It does not disclose the value, duration, or usage rights of either company’s individual agreement.

Does TollBit’s 96% referral-gap figure measure Reddit’s traffic specifically?

No. The figure compares AI-search referrals with Google Search across TollBit’s sample of 160 websites, not Reddit alone. The piece separately says Reddit reported choppy search referrals, without quantifying its own AI-search traffic loss.

What has happened in Reddit’s cases against Anthropic and Perplexity AI?

The piece identifies allegations: Anthropic allegedly accessed Reddit more than 100,000 times after saying it had stopped, and Reddit sued Perplexity AI and scraping firms over alleged DMCA violations. It provides no ruling, settlement, or other outcome.

Do Reddit contributors or moderators receive a share of AI-licensing revenue?

The piece does not describe any revenue-sharing arrangement for users or moderators. It treats moderators’ judgments as a key part of the archive’s value, but does not specify compensation tied to licensing.

AI answer interfaces and the referral gap

MeasureEarlier figureLater or comparison figureScope
Search versus social referral traffic34.8% from search by 201725.6% from social by 2017Referral-traffic comparison
AI-search referrals versus Google SearchGoogle Search baseline96% less referral traffic from AI search enginesTollBit sample of 160 websites
News searches with no click-through56% in May 202469% in May 2025Similarweb measurement after Google introduced AI Overviews

Ten percent looks small beside Reddit’s advertising business. Yet the licensing line captures contracted extraction rights alone. Reddit Answers, controlled crawler access, advertiser-linked retrieval, and the moderators who preserve source quality remain outside that figure. The 10% revenue line is the first visible quote in the archive’s order book.