/
Navigation
Chronicles
Browse all articles
Explore
Semantic exploration
Research
Entity momentum
Nexus
Correlations & relationships
Story Arc
Topic evolution
Drift Map
Semantic trajectory animation
Posts
Analysis & commentary
Pulse API
Tech news intelligence API
Browse
Entities
Companies, people, products, technologies
Domains
Browse by publication source
Handles
Browse by social media handle
Detection
Concept Search
Semantic similarity search
High Impact Stories
Top coverage by position
Sentiment Analysis
Positive/negative coverage
Anomaly Detection
Unusual coverage patterns
Analysis
Rivalry Report
Compare two entities head-to-head
Semantic Pivots
Narrative discontinuities
Crisis Response
Event recovery patterns
Connected
Search: /
Command: ⌘K
Embeddings: large
TEXXR

Chronicles

The story behind the story

← → days · ↑ ↓ browse · Enter similar · o open

OpenAI open sources Whisper, an automatic speech recognition system trained on 680K hours of “multilingual and multitask supervised data” from the web

TechCrunch Kyle Wiggers

Context & Ripple Effects

OpenAI is giving away the weights of Whisper, an automatic speech recognition system trained on 680K hours of multilingual, multitask supervised data scraped from the web — an unusually open move for a lab that would later keep its flagship models closed. A New Yorker profile of the software notes it transcribes more than 90 languages and beats humans on some of them.

The release matters less as a product than as plumbing: sources later reported OpenAI used Whisper itself to transcribe over a million hours of YouTube video as training text for GPT-4, making the model both a giveaway and an internal data-harvesting tool.

First-order effects

  • Developers get production-grade multilingual transcription for free overnight, undercutting paid speech-to-text APIs and letting anyone run ASR locally without sending audio to a vendor.
  • OpenAI gains a reusable internal pipeline: the same model it open-sources becomes the transcription engine behind its own training-data collection, per the reported YouTube harvesting effort.

Second-order effects

  • Rivals are forced to compete on coverage rather than price — Meta answers with Omnilingual ASR spanning 1,600+ languages against Whisper's 99, turning language count into the new benchmark.
  • Real-world deployments expose the cost of the free tier: engineers and researchers report Whisper hallucinating entire sentences, including racial commentary and invented medical treatments, in high-stakes uses like medical appointments.

Third-order effects

  • Basic speech-to-text commoditizes into open infrastructure, pushing differentiation up the stack toward reasoning-heavy voice agents — visible in OpenAI's later API launch of GPT-Realtime-Whisper alongside GPT-5-class realtime models.
  • The pattern — open weights paired with proprietary control planes and API monetization — becomes the template for how frontier labs release commodity capabilities while reserving value in hosted services.

The trend: Speech recognition is collapsing into free open-weight infrastructure, with labs like OpenAI and Meta competing instead on language coverage and folding transcription back into paid multimodal APIs.

Discussion

  • @gdb Greg Brockman on x
    Whisper, a neural net that approaches human level robustness and accuracy on English speech recognition. Attached is transcriptions of the same voicemail with iOS vs Whisper. Available today as open-source: https://openai.com/... https://twitter.com/...
  • @rgblong Robert Long on x
    is there going to be a step-change in writing vs transcribing behavior soon? massive leap in convenience of transcription once you don't have to talk in a halting way and the accuracy is sufficiently high speech-to-text may become the default way of getting ideas down https://twi…
  • @joannejang Joanne Jang on x
    this is huge — in addition to performing super well on accents, it also aces language switching too (like Hinglish)! https://twitter.com/...
  • @perpetualmaniac Zach Vorhies on x
    OpenAI just solved speech 2 text. According to the demo this AI is able to comprehend a fast voice in a popular commercial from the 90's. Also a thick scottish accent. The paper claims it can understand all kinds of accents due large training size. https://openai.com/...
  • @pytorch @pytorch on x
    Exciting! Clean ASR code and PyTorch models from @OpenAI. https://twitter.com/...
  • @iscienceluvr Tanishq Mathew Abraham on x
    Today, @OpenAI announced Whisper, an automatic speech recognition model. Plus, they released it open-source! Blog post → https://openai.com/... Research paper → https://cdn.openai.com/... Open-source code and models → https://github.com/... Quick thread about it (1/10) ↓ https://…
  • @marcelsalathe @marcelsalathe on x
    Impressive results! Surprise announcement: AI by OpenAI is now open. Thanks #stablediffusion 😉 https://twitter.com/...
  • @_dave__white_ Dave White on x
    permissionlessness has been one of the major accelerants of crypto, fueling innovation and enabling composability. stable diffusion has given us a taste of what it looks like in AI in recent months. today, @OpenAI released their latest model fully open source. buckle up. https://…
  • @emostaque Emad on x
    Bravo 👏 Speech2Image anyone? https://twitter.com/...
  • @openai @openai on x
    We've trained a neural net called Whisper that approaches human-level robustness and accuracy on English speech recognition. It performs well even on diverse accents and technical language. Whisper is open source for all to use. https://t.co/ueVywYPEkK
  • @deliprao Delip Rao on x
    Love this new openness from Open AI. More of that please! haven't done WER on custom corpora, but the paper suggests significant improvements over wav2vec 2.0 large on benchmarks. That's a big deal. https://twitter.com/...
  • @sama Sam Altman on x
    near human-level speech recognition, open-sourced: https://openai.com/... (check out the examples, i find them difficult)
  • @josourcing Nicole Miller on x
    @Techmeme @Kyle_L_Wiggers In its response to the USPTOffice, OpenAI acknowledges that the vast majority of content posted online is protected by U.S. copyright laws. https://www.uspto.gov/... They're using it as a 0-cost asset and selling it anyway. https://twitter.com/...