/
Navigation
Chronicles
Browse all articles
Explore
Semantic exploration
Research
Entity momentum
Nexus
Correlations & relationships
Story Arc
Topic evolution
Drift Map
Semantic trajectory animation
Posts
Analysis & commentary
Pulse API
Tech news intelligence API
Browse
Entities
Companies, people, products, technologies
Domains
Browse by publication source
Handles
Browse by social media handle
Detection
Concept Search
Semantic similarity search
High Impact Stories
Top coverage by position
Sentiment Analysis
Positive/negative coverage
Anomaly Detection
Unusual coverage patterns
Analysis
Rivalry Report
Compare two entities head-to-head
Semantic Pivots
Narrative discontinuities
Crisis Response
Event recovery patterns
Connected
Search: /
Command: ⌘K
Embeddings: large
TEXXR

Chronicles

The story behind the story

days · browse · Enter similar · o open

Sources: OpenAI transcribed 1M+ hours of YouTube videos through Whisper and used the text to train GPT-4; Google also transcribed YouTube videos to harvest text

OpenAI, Google and Meta ignored corporate policies, altered their own rules and discussed skirting copyright law …

New York Times

Context & Ripple Effects

The report lands amid a widening dispute over whether publicly accessible platform content can be repurposed as model-training input. OpenAI had reportedly considered using public YouTube transcripts for GPT-5, while YouTube's CEO said training Sora on YouTube videos would violate the platform's terms.

It also puts Google’s role under sharper scrutiny: related coverage says the company broadened its privacy policy to cover AI training on publicly available content. The core issue is not simply access to video, but whether transcription turns creator material into a usable training corpus without a negotiated license.

First-order effects

  • OpenAI and Google face heightened legal, contractual and reputational exposure over the reported use of transcribed YouTube material; YouTube creators’ work becomes central evidence in the dispute over training-data permission.
  • YouTube must reconcile its stated creator-contract protections with allegations that Google harvested text from videos, particularly after its CEO’s warning that Sora training on YouTube content would breach its terms.

Second-order effects

  • AI developers and platforms have greater incentive to document data provenance and distinguish platform access from permission to train, raising the operational value of governed, licensable corpora.
  • Creators and platforms gain leverage to seek compensation or tighter controls over training use; this is consistent with later reports of companies paying creators for unpublished video access.

Third-order effects

  • If transcription of publicly available media is treated as a separate training-data pathway, copyright, terms-of-service and privacy rules will increasingly be tested at the transformed-text layer rather than only at the original file layer.
  • The industry may move from broad web-scale collection toward more formal data-rights markets and auditable corpus governance, though the legal boundary between public availability and training permission remains unsettled.

The trend: Generative-AI developers are shifting from treating public content as abundant input toward negotiating, governing and defending access to high-quality training data.

Discussion

  • @wxdylan.bsky.social Dylan Lusk on bluesky
    So GPT-4 has the speech and mannerisms of a 10 year old, got it.  [embedded post]
  • @wavesblog Simonetta Vezzoso on x
    “Photobucket CEO Leonard says he is on solid legal ground, citing an update to the company's terms of service in October that grants it the “unrestricted right” to sell any uploaded content for the purpose of training AI systems” 🙈
  • @brij @brij on x
    Given how easy it is to run @farcaster_xyz or @joinmastodon clients I think chatgpt wrapper apps should run their own content farming game on the side. Just look at the economics of training data: $1 to $2 per image $2 to $4 per short-form video $100 to $300 per hour of longer...
  • @ppopiel Pawel Popiel on x
    Data Gold Rush: “companies are generally willing to pay $1 to $2 per image, $2 to $4 per short-form video and $100 to $300 per hour of longer films. The market rate for text is $0.001 per word. Images of nudity go for $5 to $7” https://www.reuters.com/...
  • @jacordova1961 José Antonio Córdova on x
    Photobucket image storage company, told Reuters is in talks with multiple tech companies to license its 13 billion photos and videos to be used to train generative AI models, as many other similar arrangements exist worldwide - https://www.reuters.com/...
  • @reuters @reuters on x
    Photobucket CEO Ted Leonard told Reuters he is in talks with multiple tech companies to license the company's 13 billion photos and videos to be used to train generative AI models that can produce new content in response to text prompts https://www.reuters.com/...
  • r/aiwars r on reddit
    Inside Big Tech's underground race to buy AI training data
  • r/mlscaling r on reddit
    “Inside Big Tech's underground race to buy AI training data” (even Photobucket's archives are now worth something due to data scaling)
  • r/technews r on reddit
    Inside Big Tech's underground race to buy AI training data
  • @daveleebbg Dave Lee on threads
    I can see how one million scraped YouTube videos slipped Mira Murati's mind
  • @jonathangarelick Jonathan Garelick on threads
    Common rhetorical questions: - Is the pope Catholic?  - Does a bear sh*t in the woods?  - Does OpenAI use training data from YouTube?
  • @joannastern Joanna Stern on threads
    NYTimes reports that OpenAI scraped YouTube, transcribed videos and used it to train GPT 4.  Reminder: a few weeks ago OpenAI CTO Mira Murati told me in a video interview that she didn't know if YouTube videos were used to train Sora. https://www.nytimes.com/...
  • @jonkeeganstories Jon Keegan on threads
    Fantastic story about how Big Tech companies broke their own policies (and their competitors') in the insatiable quest for more content to train their AI systems. https://www.nytimes.com/...
  • @nevillehobson Neville Hobson on threads
    It looks like anything online that's accessible is fair game, whether private or public.  This surely is IP theft, violating copyright laws, and corporate rule-breaking by OpenAI, Google, and Meta, on an industrial scale!
  • @crumbler Casey Newton on threads
    Great story about all the underhanded ways tech companies are obtaining data to train LLMs without having to (gasp) pay anyone for it https://www.nytimes.com/...
  • @glenngabe Glenn Gabe on threads
    Transcribing YouTube, 1M hours of it -> OpenAI transcribed 1M+ hours of YouTube videos through Whisper and used the text to train GPT-4; Google also transcribed YouTube videos to harvest text “Some OpenAI employees discussed how such a move might go against YouTube's rules, three…
  • @dangillmor@mastodon.social Dan Gillmor on mastodon
    Copyright is a deeply flawed weapon to use against the generative “AI” companies.  But the tech industry's arrogant vacuuming up of everything in sight — ingesting what others created so it can re-sell it to us — is moving us quickly toward the worst possible Internet, a pay-per-…
  • @Dhmspector@mastodon.social Dave Spector on mastodon
    Remember kids, when Google, Meta, OpenAI and their buddies talk about #AI it's just another word for industrial scale IP theft and copyright violations.  —  They'll steal everything your create, then come after you when you try to defend your rights to your own work.  —  https://…
  • @florian4gamers Florian Mueller on x
    Someone's conjuring up a YouTuber class action.
  • @mikeisaac Rat King on x
    that includes OpenAI, which transcribed more than 1million hours of YouTube video to train GPT-4 — which is, of course, against the terms of service of the data holders and without the knowledge of the creators but as this piece shows, all the companies are taking this route [ima…
  • @mikeisaac Rat King on x
    every company training artificial intelligence models realizes their problem is finding enough data across the internet to make their products live up to their sky-high future expectations By @CadeMetz @ceciliakang @sheeraf @stuartathompson @nicoagrant https://www.nytimes.com/...…
  • @jimprosser Jim Prosser on x
    Only 17 paragraphs down, in one sentence, is the reader reminded that the NYT is suing the subject of their article (OpenAI) over substantially related issues. Feels like this merits much more disclosure higher up. https://www.nytimes.com/...
  • @dm_cooper Danielle Miriam Cooper on x
    Poignant example of how high quality data, aka “carefully written and edited by professionals” is the most valuable for AI: “At Meta...managers, lawyers and engineers last year discussed buying the publishing house Simon & Schuster to procure long works.” https://www.nytimes.com/…
  • @dschatsky David Schatsky on x
    In May, Sam Altman, the chief executive of OpenAI, acknowledged that A.I. companies would use up all viable data on the internet. “That will run out,” he said in a speech at a tech conference. How Tech Giants Cut Corners to Harvest Data for A.I. https://www.nytimes.com/...
  • @rosenzweigjane Jane Rosenzweig on x
    “Tech companies could run through the high-quality data on the internet as soon as 2026, according to Epoch, a research institute. The companies are using the data faster than it is being produced.” https://www.nytimes.com/...
  • @jcpxdesigns @jcpxdesigns on x
    The New York Times confirmed what many of us assumed, companies making AI tools stole people's data to train their models and intentionally violated artists' copyrights because it would “take too long” to negotiate fair payment. https://www.nytimes.com/...
  • @markandrejevic Mark Andrejevic on x
    It's not just well-known authors or artists feeding the databases. It's anyone who has posted anything online. These systems represent the capture of our collective cultural and social production — they should be publicly owned and controlled. https://www.nytimes.com/...
  • @_leobriceno Leo Briceno on x
    I see a congressional hearing in the cards. Very fascinating article by NYT this morning. https://www.nytimes.com/... [image]
  • @gavatron @gavatron on x
    From today's @nytimes: AI companies need data to scale their products, and they're running out of open data. So copyright be dammed. https://www.nytimes.com/... One of many 👀👀👀 gems: [image]
  • r/technology r on reddit
    OpenAI transcribed over a million hours of YouTube videos to train GPT-4
  • r/ArtistHate r on reddit
    How Tech Giants Cut Corners to Harvest Data for A.I.
  • r/privacy r on reddit
    How Tech Giants Cut Corners to Harvest Data for A.I. (Gift Article)
  • r/mlscaling r on reddit
    OpenAI transcribed 1M+ hours of YouTube videos through Whisper and used the text to train GPT-4; Google also transcribed YouTube videos to harvest text
  • @drewharwell Drew Harwell on threads
    “What is the end goal here?” one member of the privacy team asked in an internal message.  “How broad are we going?”
  • @carnage4life Dare Obasanjo on x
    Google thinks OpenAI crawling YouTube to train its AI is “unauthorized scraping” even though that's exactly how Google also trains its AI (crawling the web & using YouTube videos) is a bit hypocritical. [image]
  • r/technology r on reddit
    OpenAI and Google reportedly used transcriptions of YouTube videos to train their AI models
  • @mgsiegler M.G. Siegler on threads
    Absolutely no one innovates like Meta when it comes to new and exciting ways to garner bad publicity.  Buying a book publisher to train AI data on its works.  Is Operation: Kick Puppies next?
  • @fletcher0xff@mastodon.online Ashley Fletcher on mastodon
    AI models are expected to run out of high-quality data by 2026, at which point there are plans to train on “synthetic” data created by AI instead of humans.  #GarbageInGarbageOut  —  https://www.nytimes.com/...
  • @dead.place Billie on bluesky
    theres something so evil about buying a storied publication just to feed its ‘content’ to the ai training gods [embedded post]
  • @puzzledpeaces.bsky.social Kathrin on bluesky
    Controversial opinion: if a company is so big they can just budget for consequences, then they are effectively immune to those consequences and must be stopped as a matter of course.  [embedded post]
  • @mikeisaac Rat King on x
    meta had these same sorts of discussions, including the idea of acquiring the publisher Simon and Schuster to scan their vast catalog of books But they all are thinking the same: doing the many, many deals this would require to not upset copyright holders would take too long. [im…