Human audio transcribers are growing in demand as AI transcription tech still struggles to parse speech from multiple speakers and grasp context
Steve Lohr / New York Times : Tweets: @nytimestech Tweets: @nytimestech : Human transcribers would seem an endangered species. Yet one business has thrived, despite the arrival of automated services and advancing A.I. technology. https://www.nytimes.com/...
Context & Ripple Effects
This 2020 piece is the early data point in a longer arc the corpus keeps returning to: automated services like Otter.ai kept adding workflow features such as meeting summaries (meeting summaries and a home feed), and OpenAI's Whisper later claimed human-level accuracy in some of its 90-plus languages (Whisper's multilingual accuracy). Yet the endpoint of that arc is not full replacement — years later, court reporters are still considered crucial for capturing gestures and working through noise, amid a worker shortage (court reporting's persistence).
First-order effects
- Transcription businesses handling multi-speaker and context-heavy audio keep pricing power over the exact segments where automated services remain least accurate.
- Automated providers like Otter.ai compete for the routine single-speaker volume instead, pushing the two tiers toward different customers rather than head-on.
Second-order effects
- Hybrid workflows emerge as the compromise: AI generates call transcripts and suggestions while workers clean up behind it, as the AT&T call center case shows — raising the question of whether humans are training their own replacements.
- Court reporting's existing worker shortage tightens further if transcription firms and courts draw from the same pool of people skilled at parsing messy, multi-party audio.
Third-order effects
- If the pattern holds, speech AI bifurcates the profession rather than eliminating it: machines absorb bulk volume while humans concentrate in verification, legal, and hard-audio work — a structure already visible in court reporting years later.
- Accuracy claims like Whisper's set customer expectations faster than real-world multi-speaker performance delivers, making 'human-in-the-loop' a durable selling point rather than a transitional one.
The trend: Speech recognition keeps automating the bulk of transcription, but consistently cedes the hardest audio — multiple speakers, noise, context — to humans, turning the profession into a verification layer atop machine output.