/
Navigation
Chronicles
Browse all articles
Explore
Semantic exploration
Research
Entity momentum
Nexus
Correlations & relationships
Story Arc
Topic evolution
Drift Map
Semantic trajectory animation
Posts
Analysis & commentary
Pulse API
Tech news intelligence API
Browse
Entities
Companies, people, products, technologies
Domains
Browse by publication source
Handles
Browse by social media handle
Detection
Concept Search
Semantic similarity search
High Impact Stories
Top coverage by position
Sentiment Analysis
Positive/negative coverage
Anomaly Detection
Unusual coverage patterns
Analysis
Rivalry Report
Compare two entities head-to-head
Semantic Pivots
Narrative discontinuities
Crisis Response
Event recovery patterns
Connected
Search: /
Command: ⌘K
Embeddings: large
TEXXR

Chronicles

The story behind the story

days · browse · Enter similar · o open

YouTube's CEO says that OpenAI training Sora with YouTube videos would violate YouTube's ToS, and Google adheres to YouTube's creator contracts to train Gemini

depending on how it trains its Sora video tool Matthias Bastian / The Decoder : YouTube CEO's warning to OpenAI over Sora training data could backfire spectacularly Cecily Mauran / Mashable : OpenAI's Sora just dropped a trippy music video to fan the AI hype flames Rich Ord / WebProNews : YouTube CEO Thows Cold Water On OpenAI Scraping Video Data Emma Roth / The Verge : OpenAI training Sora on YouTube videos would violate the platform's rules. Threads: Casey Newton / @crumbler : Honestly I dare Google to sue OpenAI for training an LLM on other people's stuff X: Michael Frank / @mfrankdude : It's not just a violation of Youtube's ToS. Scraping is a ToS violation on nearly all major platforms. @nealmohan is just the first to mention it in awhile. @alexjc wrote a python library to parse ToS verbiage and found most platforms prohibit scraping https://github.com/... @emostaque : YouTube should release a dataset of the 6m+ CC licensed videos there though [image] @modestproposal1 : Service built on copyrighted material concerned about potential TOS violation 😀 Paul Bannister / @pbannist : OMG this is priceless hypocrisy of the highest level. How it can even be said with a straight face is baffling. @emostaque : Think we were still the only company that did opt outs, took out Books3 etc, licensed music, didn't torrent movies Was harder but right way to do it Tbh future datasets will all be synthetic though Zach Edwards / @thezedwards : If Meta couldn't keep Bright Data from scraping their publicly available content, what chance does G have of keeping OpenAI from doing the same to Youtube public content? The precedent around this is getting quite clear imnlo... 🖖🏻 @brij : YouTube has come full circle. I read a funny anecdote in Mark Bergen's book about YouTube's story, where the previous CEO, Susan Wojcicki, apparently played hardball with Rupert Murdoch over a business deal Her argument - “all right, we just won't pay you. how does that sound?”... Rob Leathern / @robleathern : @alexeheath Meanwhile as you wrote “Relatedly, I heard about an internal policy put in place last year that says only AI-related efforts get access to new compute power inside Google, which means that teams running other systems have to make do with what they already have. This has,... Alessandro Perilli / @giano : The use of “if” is almost comical. Where else would OpenAI (or anybody else) find a corpus of videos for AI training as large and diverse and labelled except on YouTube? @aifray : We may see a first here: the first time for Google to find itself on that side of the debate over how online services use media content. For hell to really freeze over, they'd have to bring a copyright infringement lawsuit. @emostaque : Think it's pretty broadly assumed Whisper is trained on a massive YouTube dump with that feeding GPT-4, some interesting things in there Discovery would be unpleasant for many generative AI companies, folk trained on Hollywood movies, music downloaded & more Not stability ofc @amuse : @emilychangtv ... I think it is completely reasonable for an AI to have access to the same data that any individual has access to... Bilawal Sidhu / @bilawalsidhu : @emilychangtv ... Not unlike Reddit or Quora wanting to be paid by companies to train on their user generated content. Barry Schwartz / @rustybrick : Google to publishers - we can use your content to train our search engines and AI Google to OpenAI - you cannot use YouTube to train your AI Emily Chang / @emilychangtv : YouTube CEO @nealmohan tells me exclusively — if OpenAI is using YouTube videos to train Sora, that would be a “clear violation” of their policies (with @daveyalba) https://www.bloomberg.com/... [image] Forums: r/OpenAI : YouTube Says OpenAI Training Sora With Its Videos Would Break Rules r/ArtistHate : YouTube Says OpenAI Training Sora With its Videos Would Break the Rules See also Mediagazer

Bloomberg

Context & Ripple Effects

The warning places video-model training inside the same permissions dispute already sharpened by the New York Times’ copyright case against OpenAI and Microsoft and OpenAI’s defense of training as fair use with an opt-out. YouTube’s distinction is consequential: it says Google trains Gemini under creator contracts, while treating third-party collection of YouTube video differently.

Related coverage shows the issue extending from training inputs to generated outputs: later reporting documented Sora clips depicting well-known copyrighted characters and negotiations over licensing Hollywood content for video-generation tools.

First-order effects

  • OpenAI faces a clear platform-contract risk if Sora’s training process uses YouTube videos in a way YouTube considers unauthorized, while YouTube publicly differentiates Google’s Gemini training from outside access.
  • Creators gain a stated basis to scrutinize whether their YouTube agreements cover Google’s AI use and whether other AI developers have permission to use the same videos.

Second-order effects

  • Video-model developers may need to rely more heavily on licensed, directly supplied, or otherwise permissioned footage, increasing the strategic value of rights-holder deals.
  • Google’s contractual-access claim raises competitive pressure on rivals: data access, not only model capability, becomes a differentiator for video AI.

Third-order effects

  • The dispute points toward a split between openly accessible content and content that is contractually usable for AI training; the boundary will depend on platform terms, creator agreements, and legal outcomes.
  • If rights holders continue to resist unlicensed training, video AI commercialization is likely to move toward negotiated data rights and more formal provenance controls rather than treating public availability as permission.

The trend: Generative-video competition is increasingly becoming a contest over enforceable content rights and privileged training-data access.

Discussion

  • @crumbler Casey Newton on threads
    Honestly I dare Google to sue OpenAI for training an LLM on other people's stuff
  • @mfrankdude Michael Frank on x
    It's not just a violation of Youtube's ToS. Scraping is a ToS violation on nearly all major platforms. @nealmohan is just the first to mention it in awhile. @alexjc wrote a python library to parse ToS verbiage and found most platforms prohibit scraping https://github.com/...
  • @emostaque @emostaque on x
    YouTube should release a dataset of the 6m+ CC licensed videos there though [image]
  • @emilychangtv Emily Chang on x
    YouTube CEO @nealmohan tells me exclusively — if OpenAI is using YouTube videos to train Sora, that would be a “clear violation” of their policies (with @daveyalba) https://www.bloomberg.com/... [image]
  • @modestproposal1 @modestproposal1 on x
    Service built on copyrighted material concerned about potential TOS violation 😀
  • @pbannist Paul Bannister on x
    OMG this is priceless hypocrisy of the highest level. How it can even be said with a straight face is baffling.
  • @rustybrick Barry Schwartz on x
    Google to publishers - we can use your content to train our search engines and AI Google to OpenAI - you cannot use YouTube to train your AI
  • @emostaque @emostaque on x
    Think we were still the only company that did opt outs, took out Books3 etc, licensed music, didn't torrent movies Was harder but right way to do it Tbh future datasets will all be synthetic though
  • @thezedwards Zach Edwards on x
    If Meta couldn't keep Bright Data from scraping their publicly available content, what chance does G have of keeping OpenAI from doing the same to Youtube public content? The precedent around this is getting quite clear imnlo... 🖖🏻
  • @brij @brij on x
    YouTube has come full circle. I read a funny anecdote in Mark Bergen's book about YouTube's story, where the previous CEO, Susan Wojcicki, apparently played hardball with Rupert Murdoch over a business deal Her argument - “all right, we just won't pay you. how does that sound?”..…
  • @robleathern Rob Leathern on x
    @alexeheath Meanwhile as you wrote “Relatedly, I heard about an internal policy put in place last year that says only AI-related efforts get access to new compute power inside Google, which means that teams running other systems have to make do with what they already have. This h…
  • @giano Alessandro Perilli on x
    The use of “if” is almost comical. Where else would OpenAI (or anybody else) find a corpus of videos for AI training as large and diverse and labelled except on YouTube?
  • @aifray @aifray on x
    We may see a first here: the first time for Google to find itself on that side of the debate over how online services use media content. For hell to really freeze over, they'd have to bring a copyright infringement lawsuit.
  • @emostaque @emostaque on x
    Think it's pretty broadly assumed Whisper is trained on a massive YouTube dump with that feeding GPT-4, some interesting things in there Discovery would be unpleasant for many generative AI companies, folk trained on Hollywood movies, music downloaded & more Not stability ofc
  • @amuse @amuse on x
    @emilychangtv ... I think it is completely reasonable for an AI to have access to the same data that any individual has access to...
  • @bilawalsidhu Bilawal Sidhu on x
    @emilychangtv ... Not unlike Reddit or Quora wanting to be paid by companies to train on their user generated content.
  • r/OpenAI r on reddit
    YouTube Says OpenAI Training Sora With Its Videos Would Break Rules
  • r/ArtistHate r on reddit
    YouTube Says OpenAI Training Sora With its Videos Would Break the Rules