Sources: Nvidia scraped sources like Netflix and YouTube to train an unreleased foundational model; concerned staff were told they had full clearance to do so
Nvidia scraped videos from Youtube and several other sources to compile training data for its AI products, internal Slack chats …
404 MediaSamantha Cole
Context & Ripple Effects
This report extends earlier coverage that Nvidia was among companies whose AI training data included YouTube transcripts, placing the company in a widening dispute over whether publicly reachable media is available for model development. The new allegation concerns video collection from multiple services and an internal assertion of clearance, rather than only the provenance of a dataset.
The issue quickly moved from dataset scrutiny toward legal exposure: related coverage records a creator’s proposed class action accusing Nvidia of unauthorized scraping. That makes the reported internal authorization process consequential, not merely a technical data-sourcing detail.
First-order effects
Nvidia faces sharper scrutiny of the provenance and authorization behind training material for the unreleased model; employees who raised concerns may have their escalation and clearance processes examined.
YouTube, Netflix, and affected creators gain a more concrete basis to challenge the use of their video material, while Nvidia’s reported position is that staff had authorization.
Platforms and rights holders may have greater incentive to tighten access terms or pursue claims, raising the operational value of licensed, auditable training data.
Third-order effects
If challenges to web-scale media collection continue, the boundary between public availability and permission to train models will become a central constraint on frontier-model development.
The likely structural shift is toward training-data governance as a competitive capability: firms with clearer rights, records, and supplier relationships may face less legal and reputational friction.
The trend: This is one data point in the shift from treating online content as broadly usable AI input to treating provenance and permission as core model-development constraints.
Not surprising to see NVIDIA doing this - practically the industry standard right now - but interesting to see details of what they're collecting and why: “Movies are actually a good source of data to get gaming-like 3D consistency and fictional content but much higher quality”
A few weeks ago there was a big response to a story about companies training just on captions scraped from YouTube - captions only, not the video. This NVIDIA story involves the full video content. https://x.com/...
Employees who raised questions about ethical and legal issues involved in scraping were told this was an “executive decision” and that they had “umbrella approval” to grab whatever they could https://www.404media.co/... [image]
This is a fascinating look into how a tech giant attempts to stay competitive in the AI industry: by gobbling up as much data, including copyrighted content, as it can, as fast as it can https://www.404media.co/...
Internal Nvidia emails, conversations and documents leaked to 404 Media show how the company built a yet-to-be released video foundation model by scraping copyrighted content and academic datasets https://www.404media.co/...
This chart shows Nvidia had compiled at least 38.5 million video URLs to download. Later, CEO Jensen Huang was being given updates about their progress: “Great update,” he said. [image]
The leak also shows Nvidia scraping Netflix to make its model. Netflix says it does not have a deal with Nvidia, and that it does not allow scraping under its terms of service https://www.404media.co/... [image]