Report: Google changed its privacy policy in June 2022 to more broadly cover its use of publicly available content, including Google Docs, to train AI models
According to The New York Times, the companies may have violated YouTube creators' copyrights. — OpenAI and Google trained …
Context & Ripple Effects
The report places Google’s model-data practices alongside allegations that both Google and OpenAI drew training text from YouTube material, including the reported transcription of more than a million hours of YouTube video. It sharpens the distinction between content that is publicly reachable and content whose creators have granted a clear training right.
It also extends an earlier tension around YouTube’s exploration of AI licensing with UMG: platforms can seek negotiated rights for some catalogs while relying on broad platform terms for other material. The practical question is whether policy language provides durable permission for AI training, particularly where copyright claims remain unresolved.
First-order effects
- Google faces closer scrutiny of the scope and timing of consent embedded in its privacy policy, especially for publicly available material associated with Google Docs and YouTube.
- Creators whose YouTube works may have supplied training data gain a more concrete basis to question whether platform access and copyright permission were treated as equivalent.
Second-order effects
- YouTube and other content platforms face pressure to make AI-training terms, creator controls, and any licensing pathways more explicit rather than leaving them dispersed across general privacy policies.
- Model developers may have to weigh the legal and reputational cost of broad web-derived corpora against more traceable licensed or permissioned data sources.
Third-order effects
- If disputes continue to center on terms changes rather than solely on copying, control of AI training data will increasingly depend on how platforms define and update user consent—an important shift toward explicitly AI-oriented terms.
- The industry could split between broad-terms data collection and negotiated data licensing, with courts and policy makers determining how far contractual consent can settle copyright and creator-compensation questions.
The trend: This is one data point in AI training moving from indiscriminate access to content toward contested, governed rights over the data used to build models.