/
Navigation
Chronicles
Browse all articles
Explore
Semantic exploration
Research
Entity momentum
Nexus
Correlations & relationships
Story Arc
Topic evolution
Drift Map
Semantic trajectory animation
Posts
Analysis & commentary
Pulse API
Tech news intelligence API
Browse
Entities
Companies, people, products, technologies
Domains
Browse by publication source
Handles
Browse by social media handle
Detection
Concept Search
Semantic similarity search
High Impact Stories
Top coverage by position
Sentiment Analysis
Positive/negative coverage
Anomaly Detection
Unusual coverage patterns
Analysis
Rivalry Report
Compare two entities head-to-head
Semantic Pivots
Narrative discontinuities
Crisis Response
Event recovery patterns
Connected
Search: /
Command: ⌘K
Embeddings: large
TEXXR

Chronicles

The story behind the story

← → days · ↑ ↓ browse · Enter similar · o open

A look at OpenAI's open-source speech recognition software Whisper, which can transcribe speech in more than 90 languages, outperforming humans in some of them

OpenAI's open-source speech-transcription program—that shows us where machine learning is going. https://www.newyorker.com/... @niemanlab : “Ever since I've had tape to type up—lectures to transcribe, interviews to write down—I've dreamed of a program that would do it for me.” https://www.newyorker.com/...

New Yorker James Somers

Context & Ripple Effects

When OpenAI open-sourced Whisper on 680K hours of web audio, the pitch was accuracy at human-or-better levels across 90+ languages — and this New Yorker piece is the early-2023 portrait of what that unlocked for anyone with tape to transcribe. It reads now as the optimistic midpoint of a longer arc.

That arc has since bent twice: sources reported OpenAI used Whisper to transcribe over a million hours of YouTube videos as GPT-4 training text, and engineers later documented that the same model hallucinates chunks of text, including racial commentary. The profile captures the tool at peak promise, before both the data-provenance questions and the reliability questions surfaced.

First-order effects

  • Journalists, researchers, and developers get production-grade multilingual transcription for free, displacing paid dictation and manual transcribing workflows overnight.

Second-order effects

  • Whisper becomes dual-use infrastructure: the same open weights that power hobbyist transcription also serve as a bulk data-harvesting pipeline for training frontier models like GPT-4.

Third-order effects

  • As voice moves into commercial APIs — OpenAI's later launch of GPT-Realtime-Whisper alongside reasoning and translate models shows transcription absorbed into a broader voice stack — the hallucination findings raise the stakes for error-tolerance standards wherever transcripts feed decisions.

The trend: Speech recognition is being commoditized from standalone product into a reusable component of larger AI pipelines, with accuracy claims and provenance questions trailing behind adoption.

Discussion

  • @michaelluo Michael Luo on x
    Journos: You thought Otter was good? James Somers has identified something better. “Whisper is basically as proficient as I am at transcription.” And he explains what comes next. “...things move very fast.” https://www.newyorker.com/...
  • @pstasiatech Paul Triolo on x
    Whispers of A.I.'s Modular Future ChatGPT is in the spotlight, but it's Whisper—OpenAI's open-source speech-transcription program—that shows us where machine learning is going. https://www.newyorker.com/...
  • @niemanlab @niemanlab on x
    “Ever since I've had tape to type up—lectures to transcribe, interviews to write down—I've dreamed of a program that would do it for me.” https://www.newyorker.com/...
  • @johncassidy John Cassidy on x
    V. good piece. https://www.newyorker.com/...? utm_source=twitter&utm_medium= social&utm_campaign=onsite- share&utm_brand=the-new-yorker& utm_social-type=earned via @NewYorker