OpenAI open sources Whisper, an automatic speech recognition system trained on 680K hours of “multilingual and multitask supervised data” from the web
TechCrunchKyle Wiggers
Context & Ripple Effects
OpenAI is giving away the weights of Whisper, an automatic speech recognition system trained on 680K hours of multilingual, multitask supervised data scraped from the web — an unusually open move for a lab that would later keep its flagship models closed. A New Yorker profile of the software notes it transcribes more than 90 languages and beats humans on some of them.
The release matters less as a product than as plumbing: sources later reported OpenAI used Whisper itself to transcribe over a million hours of YouTube video as training text for GPT-4, making the model both a giveaway and an internal data-harvesting tool.
First-order effects
Developers get production-grade multilingual transcription for free overnight, undercutting paid speech-to-text APIs and letting anyone run ASR locally without sending audio to a vendor.
OpenAI gains a reusable internal pipeline: the same model it open-sources becomes the transcription engine behind its own training-data collection, per the reported YouTube harvesting effort.
Second-order effects
Rivals are forced to compete on coverage rather than price — Meta answers with Omnilingual ASR spanning 1,600+ languages against Whisper's 99, turning language count into the new benchmark.
Real-world deployments expose the cost of the free tier: engineers and researchers report Whisper hallucinating entire sentences, including racial commentary and invented medical treatments, in high-stakes uses like medical appointments.
Third-order effects
Basic speech-to-text commoditizes into open infrastructure, pushing differentiation up the stack toward reasoning-heavy voice agents — visible in OpenAI's later API launch of GPT-Realtime-Whisper alongside GPT-5-class realtime models.
The pattern — open weights paired with proprietary control planes and API monetization — becomes the template for how frontier labs release commodity capabilities while reserving value in hosted services.
The trend: Speech recognition is collapsing into free open-weight infrastructure, with labs like OpenAI and Meta competing instead on language coverage and folding transcription back into paid multimodal APIs.
Whisper, a neural net that approaches human level robustness and accuracy on English speech recognition. Attached is transcriptions of the same voicemail with iOS vs Whisper. Available today as open-source: https://openai.com/... https://twitter.com/...
is there going to be a step-change in writing vs transcribing behavior soon? massive leap in convenience of transcription once you don't have to talk in a halting way and the accuracy is sufficiently high speech-to-text may become the default way of getting ideas down https://twi…
OpenAI just solved speech 2 text. According to the demo this AI is able to comprehend a fast voice in a popular commercial from the 90's. Also a thick scottish accent. The paper claims it can understand all kinds of accents due large training size. https://openai.com/...
Today, @OpenAI announced Whisper, an automatic speech recognition model. Plus, they released it open-source! Blog post → https://openai.com/... Research paper → https://cdn.openai.com/... Open-source code and models → https://github.com/... Quick thread about it (1/10) ↓ https://…
permissionlessness has been one of the major accelerants of crypto, fueling innovation and enabling composability. stable diffusion has given us a taste of what it looks like in AI in recent months. today, @OpenAI released their latest model fully open source. buckle up. https://…
We've trained a neural net called Whisper that approaches human-level robustness and accuracy on English speech recognition. It performs well even on diverse accents and technical language. Whisper is open source for all to use. https://t.co/ueVywYPEkK
Love this new openness from Open AI. More of that please! haven't done WER on custom corpora, but the paper suggests significant improvements over wav2vec 2.0 large on benchmarks. That's a big deal. https://twitter.com/...
@Techmeme @Kyle_L_Wiggers In its response to the USPTOffice, OpenAI acknowledges that the vast majority of content posted online is protected by U.S. copyright laws. https://www.uspto.gov/... They're using it as a 0-cost asset and selling it anyway. https://twitter.com/...