Analysis: 13+ datasets used by tech companies without permission to train AI models contain 15.8M+ YouTube videos from 2M+ channels, including 1M how-to videos
www.theatlantic.com/technology/ a... Forums: r/technology : AI Is Coming for YouTube Creators | At least 15 million videos have been snatched by tech companies. See also Mediagazer
Context & Ripple Effects
This expands a pattern already visible in reporting that major AI developers trained on a dataset containing YouTube video transcripts from prominent publishers and creators. The new analysis broadens the issue from a single corpus to a dispersed training-data supply chain.
Platforms and AI companies have begun testing permissioned alternatives: YouTube introduced creator-controlled authorization for third-party AI training, while some developers reportedly paid creators for unpublished-video access.
First-order effects
- Creators and channels identified in the datasets have more concrete evidence to assess whether their work entered AI training without authorization.
- Model developers and dataset providers face sharper provenance questions over video-derived training material, particularly for instructional content that can be valuable for model capabilities.
Second-order effects
- YouTube’s opt-in authorization mechanism becomes more consequential as a way to distinguish licensed access from data gathered through third-party datasets.
- The findings strengthen incentives for AI firms to pursue negotiated creator access, as earlier reports of payments for unpublished videos suggest, rather than rely on uncertain public-web sourcing.
Third-order effects
- If repeated across modalities, training-data provenance may shift from a compliance afterthought to a core commercial input, favoring platforms and rights holders able to package and license large media libraries.
- The durable fault line is likely to be the public-data permission boundary: whether public availability is treated as sufficient for model training or requires affirmative authorization.
The trend: AI training is moving toward a contested market for auditable media rights, where creator consent and dataset provenance increasingly shape access to high-value content.