A look at Vana, which raised $20M to let users get paid to share their Reddit posts and other data to train AI models; Reddit banned Vana's subreddit
Kyle Wiggers / TechCrunch : X: @edzitron . Forums: r/technology X: Ed Zitron / @edzitron : Nuh uh uh, sorry folks. Only WE can make money off of your labor Forums: r/technology : A startup, Vana, says it wants Reddit users to get paid for training data
Context & Ripple Effects
Reddit had already moved to control the commercial use of its corpus, first by planning charges for AI-related API access while preserving developer access for Reddit-focused apps, and separately by creating a program that pays contributors for eligible on-platform activity.
Vana puts a competing ownership model in front of that strategy: individuals could attempt to license their own contributions while Reddit was pursuing large-scale data arrangements, including a reported Google data-access deal tied to search and AI training.
First-order effects
- Reddit's subreddit ban removes a key on-platform channel for Vana to organize and recruit Reddit users, constraining its effort to collect user-authorized post data.
- Users interested in Vana's payout model face a clearer conflict between sharing content individually and Reddit's control over access and distribution on its service.
Second-order effects
- The clash sharpens Reddit's incentive to route AI-data demand through its own API and licensing programs rather than allow third parties to aggregate user contributions independently.
- Other data-collection startups will need to distinguish consent-based collection from platform-governed access, especially when the source material is hosted on a platform with commercial data policies.
Third-order effects
- If platforms consistently enforce control over user-generated corpora, individual data-sharing markets may remain dependent on platform rules rather than becoming an independent supply channel for AI training data.
- The episode is part of a broader unresolved question: whether value from AI training inputs is allocated chiefly through platform licensing, direct contributor compensation, or a combination of both.
The trend: AI training data is shifting from freely harvested web content toward contested, monetized access controlled by platforms and challenged by user-compensation models.