/
Navigation
Chronicles
Browse all articles
Explore
Semantic exploration
Research
Entity momentum
Nexus
Correlations & relationships
Story Arc
Topic evolution
Drift Map
Semantic trajectory animation
Posts
Analysis & commentary
Pulse API
Tech news intelligence API
Browse
Entities
Companies, people, products, technologies
Domains
Browse by publication source
Handles
Browse by social media handle
Detection
Concept Search
Semantic similarity search
High Impact Stories
Top coverage by position
Sentiment Analysis
Positive/negative coverage
Anomaly Detection
Unusual coverage patterns
Analysis
Rivalry Report
Compare two entities head-to-head
Semantic Pivots
Narrative discontinuities
Crisis Response
Event recovery patterns
Connected
Search: /
Command: ⌘K
Embeddings: large
TEXXR

Chronicles

The story behind the story

days · browse · Enter similar · o open

Sources: Nvidia scraped sources like Netflix and YouTube to train an unreleased foundational model; concerned staff were told they had full clearance to do so

Nvidia scraped videos from Youtube and several other sources to compile training data for its AI products, internal Slack chats …

404 Media Samantha Cole

Context & Ripple Effects

This report extends earlier coverage that Nvidia was among companies whose AI training data included YouTube transcripts, placing the company in a widening dispute over whether publicly reachable media is available for model development. The new allegation concerns video collection from multiple services and an internal assertion of clearance, rather than only the provenance of a dataset.

The issue quickly moved from dataset scrutiny toward legal exposure: related coverage records a creator’s proposed class action accusing Nvidia of unauthorized scraping. That makes the reported internal authorization process consequential, not merely a technical data-sourcing detail.

First-order effects

  • Nvidia faces sharper scrutiny of the provenance and authorization behind training material for the unreleased model; employees who raised concerns may have their escalation and clearance processes examined.
  • YouTube, Netflix, and affected creators gain a more concrete basis to challenge the use of their video material, while Nvidia’s reported position is that staff had authorization.

Second-order effects

  • The allegations add pressure on AI developers to document data rights and collection methods, especially after the earlier reporting on YouTube transcripts in AI training datasets.
  • Platforms and rights holders may have greater incentive to tighten access terms or pursue claims, raising the operational value of licensed, auditable training data.

Third-order effects

  • If challenges to web-scale media collection continue, the boundary between public availability and permission to train models will become a central constraint on frontier-model development.
  • The likely structural shift is toward training-data governance as a competitive capability: firms with clearer rights, records, and supplier relationships may face less legal and reputational friction.

The trend: This is one data point in the shift from treating online content as broadly usable AI input to treating provenance and permission as core model-development constraints.

Discussion

  • @mkbhd Marques Brownlee on x
    cool cool cool cool cool cool now leaked NVIDIA slack messages discussing which YouTube channels to scrape videos from. MKBHD videos? Yeah grab those too.
  • @simonw Simon Willison on x
    Not surprising to see NVIDIA doing this - practically the industry standard right now - but interesting to see details of what they're collecting and why: “Movies are actually a good source of data to get gaming-like 3D consistency and fictional content but much higher quality”
  • @simonw Simon Willison on x
    A few weeks ago there was a big response to a story about companies training just on captions scraped from YouTube - captions only, not the video. This NVIDIA story involves the full video content. https://x.com/...
  • @samleecole Samantha Cole on x
    Employees who raised questions about ethical and legal issues involved in scraping were told this was an “executive decision” and that they had “umbrella approval” to grab whatever they could https://www.404media.co/... [image]
  • @jason_koebler Jason Koebler on x
    @MKBHD “Should we download the whole Netflix too? How would we operationalize it?” [image]
  • @jason_koebler Jason Koebler on x
    Here, Nvidia employees discuss specific YouTube channels they want to scrape, including @MKBHD's ("super high quality") [image]
  • @minimaxir Max Woolf on x
    I have to give credit for 404 Media for their continually-good sourcing.
  • @samleecole Samantha Cole on x
    This is a fascinating look into how a tech giant attempts to stay competitive in the AI industry: by gobbling up as much data, including copyrighted content, as it can, as fast as it can https://www.404media.co/...
  • @minimaxir Max Woolf on x
    I don't even know logistically how you would scrape Netflix. Torrent all the shows? https://x.com/...
  • @samleecole Samantha Cole on x
    Internal Nvidia emails, conversations and documents leaked to 404 Media show how the company built a yet-to-be released video foundation model by scraping copyrighted content and academic datasets https://www.404media.co/...
  • @jason_koebler Jason Koebler on x
    This chart shows Nvidia had compiled at least 38.5 million video URLs to download. Later, CEO Jensen Huang was being given updates about their progress: “Great update,” he said. [image]
  • @josephfcox Joseph Cox on x
    The leak also shows Nvidia scraping Netflix to make its model. Netflix says it does not have a deal with Nvidia, and that it does not allow scraping under its terms of service https://www.404media.co/... [image]
  • r/ChatGPT r on reddit
    Leaked Documents Show Nvidia Scraping ‘A Human Lifetime’ of Videos Per Day to Train AI
  • r/hardware r on reddit
    Leaked Documents Show Nvidia Scraping ‘A Human Lifetime’ of Videos Per Day to Train AI
  • r/NVDA_Stock r on reddit
    Leaked Documents Show Nvidia Scraping ‘A Human Lifetime’ of Videos Per Day to Train AI
  • r/technology r on reddit
    Leaked Documents Show Nvidia Scraping ‘A Human Lifetime’ of Videos Per Day to Train AI
  • r/ArtistHate r on reddit
    Leaked Documents Show Nvidia Scraping ‘A Human Lifetime’ of Videos Per Day to Train ML
  • r/singularity r on reddit
    Leaked Documents Show Nvidia Scraping ‘A Human Lifetime’ of Videos Per Day to Train AI