Engineers, developers, and researchers say OpenAI's transcription tool Whisper hallucinates chunks of text or even entire sentences, including racial commentary
Tech behemoth OpenAI has touted its artificial intelligence-powered transcription tool Whisper as having near “human level robustness and accuracy.”
Associated Press
Context & Ripple Effects
Whisper’s reach grew from OpenAI’s open-sourcing of the speech-recognition system to coverage emphasizing multilingual transcription performance. That history makes reported fabricated passages more consequential than ordinary transcription errors: the output can be mistaken for a record of what was said.
The concern also sits alongside reports that OpenAI used Whisper to create training text from a large volume of YouTube video, expanding the tool’s role in data processing. The immediate issue is therefore both product reliability and the treatment of machine-generated text downstream.
First-order effects
Users of Whisper, particularly those transcribing sensitive recordings such as medical appointments, must treat transcripts as drafts requiring human review when they contain unsupported or harmful language.
OpenAI’s accuracy positioning for Whisper faces sharper scrutiny because the reported failures include whole invented sentences, not just missed words or punctuation errors.
Second-order effects
Organizations embedding transcription into documentation, search, or summarization workflows may add confidence checks, audio-to-text verification, and escalation paths before generated transcripts enter records.
Competing speech-recognition providers can differentiate on auditability and error handling, while buyers may weigh those controls more heavily than headline accuracy claims.
Third-order effects
If such failures persist across speech-to-text systems, transcription will increasingly be governed as an assisted workflow rather than a reliable source-of-record layer, especially in high-consequence settings.
The episode reinforces a broader push toward operational assurance for AI: evaluation must account for unsupported content and harmful insertions, not only aggregate transcription accuracy.
The trend: AI transcription is moving from a standalone accuracy benchmark toward workflow-native deployment where provenance, review, and failure containment determine trust.
'"This [propensity of Whisper to hallucinate] seems solvable if the company is willing to prioritize it," said William Saunders, a San Francisco-based research engineer who quit OpenAI in February over concerns with the company's direction.' — *How*, William? …
Human error has the crucial quality of being human, and the liability insurance people are going to have clauses in all of their contracts soon that basically read “if you bring AI tools into this part of this process, we're out. You are on your fucking own.” — https://abcnews…
Researchers have found that an #AI powered transcription tool Whisper used in hospitals can generate false information, or “hallucinations,” which may include invented sentences, racial commentary, violent rhetoric, and fictional medical treatments. — https://apnews.com/...
I've said this before, but before the AI hype boom we already had tools that used machine learning to do stuff like ramshackle transcription. — I used Google Recorder transcription for interviews! Was pretty solid for personal purposes. — So how have we stepped *backwards* i…
This is a not a safe product and you'd be much better off hiring experienced human transcriptionists. It's an actual skill and (for now) a job humans can capably do. [image]
An #AI tool for transcribing medical appointments “is prone to making up chunks of text or even entire sentences... can include racial commentary, violent rhetoric and even imagined medical treatments.” Whoops! Almost like it was trained from Reddit... https://apnews.com/...
Researchers have found that an #AI-powered transcription tool Whisper used in hospitals can generate false information, or “hallucinations,” which may include invented sentences, racial commentary, violent rhetoric, and fictional medical treatments. https://apnews.com/...