/
Navigation
Chronicles
Browse all articles
Explore
Semantic exploration
Research
Entity momentum
Nexus
Correlations & relationships
Story Arc
Topic evolution
Drift Map
Semantic trajectory animation
Posts
Analysis & commentary
Pulse API
Tech news intelligence API
Browse
Entities
Companies, people, products, technologies
Domains
Browse by publication source
Handles
Browse by social media handle
Detection
Concept Search
Semantic similarity search
High Impact Stories
Top coverage by position
Sentiment Analysis
Positive/negative coverage
Anomaly Detection
Unusual coverage patterns
Analysis
Rivalry Report
Compare two entities head-to-head
Semantic Pivots
Narrative discontinuities
Crisis Response
Event recovery patterns
Connected
Search: /
Command: ⌘K
Embeddings: large
TEXXR

Chronicles

The story behind the story

← → days · ↑ ↓ browse · Enter similar · o open

A look at tech giants' AI training data deals; Defined.ai: some are ready to pay $1-$2 per image, $2-$4 per short video, and $100-$300 per hour of long video

Reuters

Context & Ripple Effects

Defined.ai's reported rates put a visible price schedule on training inputs, particularly video, rather than treating data as a freely available byproduct of the web. The later example of Adobe paying photographers for everyday-action video reinforces that demand can translate into structured sourcing programs.

The coverage arc broadens from licensed media to payments to creators for unpublished video, suggesting that differentiated, harder-to-replace material is becoming a procurement category for model builders.

First-order effects

  • Data providers, photographers, videographers, and other rights holders gain concrete reference points for negotiating AI-training licenses; long-form video carries the highest reported rate range.
  • Tech giants seeking video training material face a more explicit data-acquisition cost, alongside the operational work of finding, licensing, and delivering usable footage.

Second-order effects

  • Content platforms and specialist data brokers have an incentive to package footage by format, subject matter, and usage rights, since the reported rates distinguish sharply between images, short video, and long video.
  • Buyers may seek lower-cost ways to obtain comparable training value—through narrower datasets or alternative collection channels—as later coverage of smaller training datasets at Chinese AI companies illustrates.

Third-order effects

  • If these transactions persist, training data is likely to be managed less as an open-ended input and more as a commercial supply chain in which provenance, permissions, and format-specific quality affect model-development economics.
  • The market could split between premium licensed data for differentiated capabilities and lower-cost data strategies for price-sensitive model development; the corpus supports the pressure, not a settled outcome.

The trend: AI development is moving toward the commercialization of high-value training data, with video and other scarce, usable inputs acquiring explicit market prices.

Discussion

  • @wavesblog Simonetta Vezzoso on x
    “Photobucket CEO Leonard says he is on solid legal ground, citing an update to the company's terms of service in October that grants it the “unrestricted right” to sell any uploaded content for the purpose of training AI systems” 🙈
  • @brij @brij on x
    Given how easy it is to run @farcaster_xyz or @joinmastodon clients I think chatgpt wrapper apps should run their own content farming game on the side. Just look at the economics of training data: $1 to $2 per image $2 to $4 per short-form video $100 to $300 per hour of longer...
  • @ppopiel Pawel Popiel on x
    Data Gold Rush: “companies are generally willing to pay $1 to $2 per image, $2 to $4 per short-form video and $100 to $300 per hour of longer films. The market rate for text is $0.001 per word. Images of nudity go for $5 to $7” https://www.reuters.com/...
  • @jacordova1961 José Antonio Córdova on x
    Photobucket image storage company, told Reuters is in talks with multiple tech companies to license its 13 billion photos and videos to be used to train generative AI models, as many other similar arrangements exist worldwide - https://www.reuters.com/...
  • @reuters @reuters on x
    Photobucket CEO Ted Leonard told Reuters he is in talks with multiple tech companies to license the company's 13 billion photos and videos to be used to train generative AI models that can produce new content in response to text prompts https://www.reuters.com/...
  • r/aiwars r on reddit
    Inside Big Tech's underground race to buy AI training data
  • r/mlscaling r on reddit
    “Inside Big Tech's underground race to buy AI training data” (even Photobucket's archives are now worth something due to data scaling)
  • r/technews r on reddit
    Inside Big Tech's underground race to buy AI training data