A look at tech giants' AI training data deals; Defined.ai: some are ready to pay $1-$2 per image, $2-$4 per short video, and $100-$300 per hour of long video
Reuters : X: @ppopiel , @jacordova1961 , and @reuters Forums: Hacker News and r/technews X: Pawel Popiel / @ppopiel : Data Gold Rush: “companies are generally willing to pay $1 to $2 per image, $2 to $4 per short-form video and $100 to $300 per hour of longer films. The market rate for text is $0.001 per word. Images of nudity go for $5 to $7” https://www.reuters.com/... José Antonio Córdova / @jacordova1961 : Photobucket image storage company, told Reuters is in talks with multiple tech companies to license its 13 billion photos and videos to be used to train generative AI models, as many other similar arrangements exist worldwide - https://www.reuters.com/... @reuters : Photobucket CEO Ted Leonard told Reuters he is in talks with multiple tech companies to license the company's 13 billion photos and videos to be used to train generative AI models that can produce new content in response to text prompts https://www.reuters.com/... Forums: Hacker News : Big Tech's underground race to buy AI training data r/technews : Inside Big Tech's underground race to buy AI training data
Context & Ripple Effects
Training-data access was already becoming a paid product: Reddit had signed an AI-training agreement reported at about $60 million annualized, while Stack Overflow planned to charge large model developers for its question-and-answer archive. This report supplies market-level price signals for the underlying media assets.
Photobucket's talks to license its photo and video archive show that holders of large, organized libraries can turn dormant inventories into an AI-facing revenue opportunity.
First-order effects
- Defined.ai's reported rates give data owners and brokers concrete benchmarks for negotiating image, short-video, long-video, text, and sensitive-image training licenses.
- Photobucket and similarly situated archives gain evidence that their catalogues may command direct licensing fees from generative-AI developers.
Second-order effects
- Model developers seeking video capability face a clearer content-acquisition cost line item, especially for long-form footage, rather than treating training material as an undifferentiated input.
- Platforms and specialist repositories have greater incentive to package searchable, rights-managed collections, following the direction signaled by Reddit's reported AI-training agreement and Stack Overflow's planned developer access fees.
Third-order effects
- If these transactions persist, proprietary and well-documented media libraries could become a more important competitive input for model builders, alongside compute and model talent.
- A tiered market may emerge in which scarcity, format, and sensitivity determine data value; the reported premium for certain imagery suggests that not all training data will be priced alike.
The trend: Generative-AI developers are shifting from opportunistic web-scale collection toward a commercial market for licensed, structured training data.