Reddit releases a new content policy, including a ban on AI data licensees from using deleted posts or comments; Reddit expects $60M+ in 2024 licensing revenue
Context & Ripple Effects
Reddit’s policy formalizes a constraint around a licensing business that had already taken shape through a reported annualized content-training agreement worth about $60 million and a separate Google data-access and AI-training deal.
The important shift is from treating Reddit’s corpus as a bulk licensed asset to treating deletion status as a condition of continued eligibility for that asset. That makes data stewardship part of the product Reddit is selling to AI customers.
First-order effects
- AI data licensees must exclude deleted Reddit posts and comments from licensed datasets, requiring their ingestion and retention practices to reflect removals.
- Reddit gains a clearer policy basis for its projected licensing revenue while placing a user-content boundary on what its buyers can retain and use.
Second-order effects
- Licensees may need deletion-aware data pipelines and contractual clarity on how removals affect existing training corpora, raising the operational burden of using continuously updated community data.
- Other platforms pursuing AI-data licensing face pressure to define similarly explicit permissions, deletion handling, and buyer obligations rather than sell access as a one-time data transfer.
Third-order effects
- If this approach spreads, AI training-data deals will increasingly compete on provenance, permissions, and revocation handling—not simply corpus scale or access price.
- The broader market could move toward recurring, governed access to public conversation data, with platforms retaining more control over downstream model inputs; the extent depends on whether buyers accept the added compliance burden.
The trend: AI-content licensing is evolving from bulk data access into governed, revocable access to continuously changing user-generated data.