/
Navigation
Chronicles
Browse all articles
Explore
Semantic exploration
Research
Entity momentum
Nexus
Correlations & relationships
Story Arc
Topic evolution
Drift Map
Semantic trajectory animation
Posts
Analysis & commentary
Pulse API
Tech news intelligence API
Browse
Entities
Companies, people, products, technologies
Domains
Browse by publication source
Handles
Browse by social media handle
Detection
Concept Search
Semantic similarity search
High Impact Stories
Top coverage by position
Sentiment Analysis
Positive/negative coverage
Anomaly Detection
Unusual coverage patterns
Analysis
Rivalry Report
Compare two entities head-to-head
Semantic Pivots
Narrative discontinuities
Crisis Response
Event recovery patterns
Connected
Search: /
Command: ⌘K
Embeddings: large
TEXXR

Chronicles

The story behind the story

days · browse · Enter similar · o open

Analysis: 13+ datasets used by tech companies without permission to train AI models contain 15.8M+ YouTube videos from 2M+ channels, including 1M how-to videos

www.theatlantic.com/technology/ a... Forums: r/technology : AI Is Coming for YouTube Creators |  At least 15 million videos have been snatched by tech companies. See also Mediagazer

The Atlantic Alex Reisner

Context & Ripple Effects

This expands a pattern already visible in reporting that major AI developers trained on a dataset containing YouTube video transcripts from prominent publishers and creators. The new analysis broadens the issue from a single corpus to a dispersed training-data supply chain.

Platforms and AI companies have begun testing permissioned alternatives: YouTube introduced creator-controlled authorization for third-party AI training, while some developers reportedly paid creators for unpublished-video access.

First-order effects

  • Creators and channels identified in the datasets have more concrete evidence to assess whether their work entered AI training without authorization.
  • Model developers and dataset providers face sharper provenance questions over video-derived training material, particularly for instructional content that can be valuable for model capabilities.

Second-order effects

  • YouTube’s opt-in authorization mechanism becomes more consequential as a way to distinguish licensed access from data gathered through third-party datasets.
  • The findings strengthen incentives for AI firms to pursue negotiated creator access, as earlier reports of payments for unpublished videos suggest, rather than rely on uncertain public-web sourcing.

Third-order effects

  • If repeated across modalities, training-data provenance may shift from a compliance afterthought to a core commercial input, favoring platforms and rights holders able to package and license large media libraries.
  • The durable fault line is likely to be the public-data permission boundary: whether public availability is treated as sufficient for model training or requires affirmative authorization.

The trend: AI training is moving toward a contested market for auditable media rights, where creator consent and dataset provenance increasingly shape access to high-value content.

Discussion

  • @bcmerchant Brian Merchant on bluesky
    AI Watchdog is indeed an excellent resource and is, for example, how I know that my books have been ingested into the LibGen dataset and used without my consent to train commercial AI systems [embedded post]
  • @damonberes.com Damon Beres on bluesky
    I'm proud to share a new project that @alexreisner.bsky.social and I launched at @theatlantic.com today!  It's called AI Watchdog, and it's our new home for all of the investigations into training data sets, such as LibGen, Books3, and OpenSubtitles. www.theatlantic.com/category/…
  • @annakornbluh Anna Kornbluh on bluesky
    the ingenuous scholarly and amateur instructional material on youtube, and its large proportion of overall content there, is one of the 21st century's great collective achievements and its enclosure it for the profit of the tiny grifter few is very terrible [embedded post]
  • @jasonkoebler Jason Koebler on bluesky
    This is a really important resource and reporting cataloguing the scale at which AI is scraping the internet and allowing people to search these datasets  —  www.theatlantic.com/technology/ a...
  • r/technology r on reddit
    AI Is Coming for YouTube Creators |  At least 15 million videos have been snatched by tech companies.