Sources: Google is using its expansive library of YouTube videos to train Gemini and Veo 3; Google says it only uses a subset of videos for the training
Google is using its expansive library of YouTube videos to train its artificial intelligence models, including Gemini and the Veo 3 video and audio generator, CNBC has learned.
Context & Ripple Effects
The report follows Google’s May rollout of Veo 3, Imagen 4 and the Flow filmmaking tool, tying the company’s generative-video product push to a platform it already operates at global scale.
It matters because YouTube is not merely a distribution channel in this account: a selected portion of its video library is an input to Google’s multimodal model development, linking content ownership, model capability and product delivery.
First-order effects
- Google can use a subset of YouTube’s video corpus to develop Gemini and Veo 3, giving its multimodal teams an internal source of video material rather than relying solely on externally assembled datasets.
- The disclosure makes YouTube’s role in Google’s AI stack more consequential for the platform, its uploaders and users, even though Google says only a subset of videos is used.
Second-order effects
- Rivals building video-capable models face greater pressure to secure comparable multimodal training inputs or differentiate through model design, partnerships or narrower product focus.
- A platform-owned corpus can be converted into commercial model features; the subsequent Veo 3 API launch shows how video-generation capability can become a metered developer service.
Third-order effects
- If major platforms continue training models on content they host, control of large, governable media libraries may become a more durable AI advantage than model access alone.
- The pattern raises the strategic importance of corpus governance: firms will need to balance training utility with clearer controls and expectations around how hosted content is used.
The trend: Generative-video competition is shifting toward vertically integrated AI stacks in which platforms combine proprietary content inputs, multimodal models and distribution channels.