/
Navigation
Chronicles
Browse all articles
Explore
Semantic exploration
Research
Entity momentum
Nexus
Correlations & relationships
Story Arc
Topic evolution
Drift Map
Semantic trajectory animation
Posts
Analysis & commentary
Pulse API
Tech news intelligence API
Browse
Entities
Companies, people, products, technologies
Domains
Browse by publication source
Handles
Browse by social media handle
Detection
Concept Search
Semantic similarity search
High Impact Stories
Top coverage by position
Sentiment Analysis
Positive/negative coverage
Anomaly Detection
Unusual coverage patterns
Analysis
Rivalry Report
Compare two entities head-to-head
Semantic Pivots
Narrative discontinuities
Crisis Response
Event recovery patterns
Connected
Search: /
Command: ⌘K
Embeddings: large
TEXXR

Chronicles

The story behind the story

days · browse · Enter similar · o open

Sources: OpenAI transcribed 1M+ hours of YouTube videos through Whisper and used the text to train GPT-4; Google also transcribed YouTube videos to harvest text

OpenAI, Google and Meta ignored corporate policies, altered their own rules and discussed skirting copyright law …

New York Times

Context & Ripple Effects

The report follows coverage that OpenAI had considered public YouTube transcripts for GPT-5 amid warnings that high-quality training text could become scarce. It also lands after YouTube's CEO said using YouTube videos for Sora training would violate the platform's terms, sharpening the conflict over YouTube creator contracts.

Google's reported 2022 privacy-policy expansion for public content provides a related example of companies broadening the stated basis for AI-data use. Together, the coverage makes the provenance of model inputs—not simply whether data is publicly accessible—a central issue.

First-order effects

  • OpenAI and Google face immediate scrutiny over the reported transcription and use of YouTube material, while Meta is implicated by the broader allegations of internal policy changes and copyright-law avoidance discussions.
  • YouTube and its creators gain a clearer basis to press for enforcement of platform terms and greater disclosure about how uploaded material is repurposed for model training.

Second-order effects

  • Model developers relying on web-scale video or transcript corpora will face pressure to document data lineage and distinguish platform access from permission to train on derived text.
  • The dispute strengthens the commercial case for negotiated creator and publisher access, rather than treating transcription as a workaround to obtain training text.

Third-order effects

  • If rights holders and platforms consistently challenge derived training data, competitive advantage will shift toward firms with governed, auditable corpora and durable licensing relationships.
  • The boundary between public availability and authorized AI use is likely to become a core policy and contract question, potentially reshaping how platforms control downstream use of creator content.

The trend: Generative-AI builders are moving from opportunistic web-data collection toward a contested market for permissioned, traceable training corpora.

Discussion

  • @wxdylan.bsky.social Dylan Lusk on bluesky
    So GPT-4 has the speech and mannerisms of a 10 year old, got it.  [embedded post]
  • @wavesblog Simonetta Vezzoso on x
    “Photobucket CEO Leonard says he is on solid legal ground, citing an update to the company's terms of service in October that grants it the “unrestricted right” to sell any uploaded content for the purpose of training AI systems” 🙈
  • @brij @brij on x
    Given how easy it is to run @farcaster_xyz or @joinmastodon clients I think chatgpt wrapper apps should run their own content farming game on the side. Just look at the economics of training data: $1 to $2 per image $2 to $4 per short-form video $100 to $300 per hour of longer...
  • @ppopiel Pawel Popiel on x
    Data Gold Rush: “companies are generally willing to pay $1 to $2 per image, $2 to $4 per short-form video and $100 to $300 per hour of longer films. The market rate for text is $0.001 per word. Images of nudity go for $5 to $7” https://www.reuters.com/...
  • @jacordova1961 José Antonio Córdova on x
    Photobucket image storage company, told Reuters is in talks with multiple tech companies to license its 13 billion photos and videos to be used to train generative AI models, as many other similar arrangements exist worldwide - https://www.reuters.com/...
  • @reuters @reuters on x
    Photobucket CEO Ted Leonard told Reuters he is in talks with multiple tech companies to license the company's 13 billion photos and videos to be used to train generative AI models that can produce new content in response to text prompts https://www.reuters.com/...
  • r/aiwars r on reddit
    Inside Big Tech's underground race to buy AI training data
  • r/mlscaling r on reddit
    “Inside Big Tech's underground race to buy AI training data” (even Photobucket's archives are now worth something due to data scaling)
  • r/technews r on reddit
    Inside Big Tech's underground race to buy AI training data