/
Navigation
Chronicles
Browse all articles
Explore
Semantic exploration
Research
Entity momentum
Nexus
Correlations & relationships
Story Arc
Topic evolution
Drift Map
Semantic trajectory animation
Posts
Analysis & commentary
Pulse API
Tech news intelligence API
Browse
Entities
Companies, people, products, technologies
Domains
Browse by publication source
Handles
Browse by social media handle
Detection
Concept Search
Semantic similarity search
High Impact Stories
Top coverage by position
Sentiment Analysis
Positive/negative coverage
Anomaly Detection
Unusual coverage patterns
Analysis
Rivalry Report
Compare two entities head-to-head
Semantic Pivots
Narrative discontinuities
Crisis Response
Event recovery patterns
Connected
Search: /
Command: ⌘K
Embeddings: large
TEXXR

Chronicles

The story behind the story

days · browse · Enter similar · o open

Report: Google changed its privacy policy in June 2022 to more broadly cover its use of publicly available content, including Google Docs, to train AI models

According to The New York Times, the companies may have violated YouTube creators' copyrights.  —  OpenAI and Google trained …

Engadget Cheyenne MacDonald

Context & Ripple Effects

The report places Google’s model-data practices alongside allegations that both Google and OpenAI drew training text from YouTube material, including the reported transcription of more than a million hours of YouTube video. It sharpens the distinction between content that is publicly reachable and content whose creators have granted a clear training right.

It also extends an earlier tension around YouTube’s exploration of AI licensing with UMG: platforms can seek negotiated rights for some catalogs while relying on broad platform terms for other material. The practical question is whether policy language provides durable permission for AI training, particularly where copyright claims remain unresolved.

First-order effects

  • Google faces closer scrutiny of the scope and timing of consent embedded in its privacy policy, especially for publicly available material associated with Google Docs and YouTube.
  • Creators whose YouTube works may have supplied training data gain a more concrete basis to question whether platform access and copyright permission were treated as equivalent.

Second-order effects

  • YouTube and other content platforms face pressure to make AI-training terms, creator controls, and any licensing pathways more explicit rather than leaving them dispersed across general privacy policies.
  • Model developers may have to weigh the legal and reputational cost of broad web-derived corpora against more traceable licensed or permissioned data sources.

Third-order effects

  • If disputes continue to center on terms changes rather than solely on copying, control of AI training data will increasingly depend on how platforms define and update user consent—an important shift toward explicitly AI-oriented terms.
  • The industry could split between broad-terms data collection and negotiated data licensing, with courts and policy makers determining how far contractual consent can settle copyright and creator-compensation questions.

The trend: This is one data point in AI training moving from indiscriminate access to content toward contested, governed rights over the data used to build models.

Discussion

  • @drewharwell Drew Harwell on threads
    “What is the end goal here?” one member of the privacy team asked in an internal message.  “How broad are we going?”
  • @carnage4life Dare Obasanjo on x
    Google thinks OpenAI crawling YouTube to train its AI is “unauthorized scraping” even though that's exactly how Google also trains its AI (crawling the web & using YouTube videos) is a bit hypocritical. [image]
  • r/technology r on reddit
    OpenAI and Google reportedly used transcriptions of YouTube videos to train their AI models