Source and doc: Runway scraped thousands of videos from YouTube creators and brands, including Disney and VICE News, to train its Gen-3 AI video generation tool
A leaked internal document shows Runway's celebrated Gen-3 AI video generator collected thousands of YouTube videos and pirated movies for training data. X: @mkbhd , @malwarejake , @ednewtonrex , @josephfcox , @paulroetzer , @emanuelmaiberg , @emanuelmaiberg , @samleecole , @samleecole , and @samleecole X: Marques Brownlee / @mkbhd : Well well well. Runway AI video generator was trained on YouTube videos without permission, including 1600+ MKBHD videos 🫠 https://www.404media.co/... Jake Williams / @malwarejake : Most generative AI works on the misuse (theft) of data used without proper license. Courts need to step in and penalize TF out of these AI companies to send a message that this behavior won't be tolerated. Ed Newton-Rex / @ednewtonrex : 404 got access to an internal Runway spreadsheet suggesting their AI video model is trained on copyrighted YouTube videos, including from @Disney @NewYorker @Casey @MKBHD & many more. It looks like 404 have done red-teaming that backs that up. The first AI video lawsuit, when it [image] Joseph Cox / @josephfcox : The spreadsheet includes the New Yorker, VICE News, Pixar, Disney, Netflix, Sony, then influencers like Casey Neistat, Sam Kolder, Benjamin Hardman, Marques Brownlee. In a spreadsheet of potential targets by a multibillion dollar company to rip off https://www.404media.co/... [image] Paul Roetzer / @paulroetzer : My goodness. This is some excellent investigative journalism, and a major problem brewing for Runway and other AI model companies. It seems like it's only a matter of time until a landmark IP lawsuit hits that send shockwaves through the frontier model companies. Emanuel Maiberg / @emanuelmaiberg : this also is just proving what everyone basically already knows: All these AI video generators are scraping YouTube, ripping off everyone from Disney to YouTube channels operated by small teams and individuals https://www.404media.co/... Emanuel Maiberg / @emanuelmaiberg : Runway has investment from Google, and Google told us this would violate YouTube policies. Unclear how this company can continue to not comment on this https://www.404media.co/... Samantha Cole / @samleecole : Runway has previously been cagey about how it trained this model, saying it used “an in-house research team that oversees all of our training and we use curated, internal datasets.” Apparently, that means ripping off huge brands and small creators alike https://www.404media.co/... Samantha Cole / @samleecole : On the list is DEFY productions, and food YouTuber Mark Wiens; Runway produced these videos with just their names or “in the style of” prompts. Pretty damn close! https://www.404media.co/... [video] Samantha Cole / @samleecole : Runway embarked on a company-wide effort to scrape tens of thousands of YouTube videos into its training data for Gen-3 Alpha, a former employee told 404 Media. The entire document is here: https://www.404media.co/...
Context & Ripple Effects
This report adds documentary detail to an emerging dispute over whether publicly accessible video can be treated as model input without permission. It follows YouTube’s stated position that training Sora on its videos would breach its terms, in YouTube's warning over Sora training data.
The issue was already moving on two tracks: AI companies were discussing video-content licenses with Hollywood studios, while investigations identified YouTube transcripts in other companies’ datasets. The reported Runway materials make the rights question concrete for both large media owners and individual creators.
First-order effects
- Named creators, studios, and publishers gain a specific reported dataset record to review when deciding whether to seek explanations, removal, licensing, or legal remedies from Runway.
- Runway’s Gen-3 product and its Google-backed business face sharper scrutiny over training-data provenance, especially because the reported list includes both creator videos and pirated films.
Second-order effects
- Video-model developers and their investors face greater pressure to document source, permissions, and exclusions for training corpora; opaque web-scale collection becomes harder to defend commercially.
- The report strengthens the bargaining position of studios and platforms in licensing discussions, including the reported talks between AI firms and Hollywood studios, while creators have clearer reason to demand comparable treatment.
Third-order effects
- If rights holders continue to obtain dataset-level evidence, training-data disputes may shift from broad arguments about web access toward model-specific proof of copying, provenance, and authorization.
- The later Disney and NBCUniversal suit against Midjourney shows how video-generation copyright conflicts can progress from disputed data practices to direct litigation; whether that becomes a durable licensing regime remains unresolved.
The trend: Generative-video competition is increasingly constrained by a public-data permission boundary, as rights holders seek to convert training inputs into licensed, attributable assets.