Microsoft AI debuts MAI-Transcribe-2, a speech recognition model that it says beats Gemini 3.5 Transcribe and GPT-Transcribe, at $0.10/audio hour through 2026
Then it priced the thing at 10 cents per hour of audio. — That figure deserves a pause. … Thursday's early-bird price cuts that by roughly 72%.
Context & Ripple Effects
Microsoft’s April release of its first in-house MAI transcription model framed speech as part of a broader push for AI self-sufficiency. MAI-Transcribe-2 turns that effort into a price-and-performance challenge to Google and OpenAI, with Microsoft claiming leadership over their named transcription models.
Microsoft has long treated transcription as a collaboration feature, from speaker-based Teams transcription to Teams Premium’s planned AI notes. A per-audio-hour API price makes the economics of that capability more explicit for developers and high-volume users.
First-order effects
- Developers processing recorded audio can buy MAI-Transcribe-2 at $0.10 per audio hour through 2026, materially lowering the stated early-bird cost for a core speech-to-text input.
- Microsoft gains an in-house speech model it says outperforms Gemini 3.5 Transcribe and GPT-Transcribe, strengthening its ability to sell speech recognition without relying on a competing model provider.
Second-order effects
- Google and OpenAI face sharper pressure to defend transcription workloads on measurable accuracy and per-hour inference cost rather than model brand alone.
- Meeting-note, speech-analysis, and transcription-product vendors gain room to price services around workflow features and verification, because raw audio conversion represents a smaller share of their bill.
Third-order effects
- If comparable accuracy is sustained at these rates, speech recognition becomes less of a standalone premium feature and more of a commodity inference input, shifting differentiation toward the products that organize and act on transcripts.
- The move reinforces an AI-margin contest in which model providers compete by lowering inference cost while application vendors must show value beyond the underlying model call.
The trend: Speech AI is moving toward competition on cost per processed hour and benchmarked accuracy, with value migrating from transcription itself to the workflows built around it.