Microsoft AI debuts MAI-Transcribe-2, a speech recognition model that it says beats Gemini 3.5 Transcribe and GPT-Transcribe, at $0.10/audio hour through 2026
Then it priced the thing at 10 cents per hour of audio. — That figure deserves a pause. … Thursday's early-bird price cuts that by roughly 72%.
Context & Ripple Effects
Microsoft AI has been building a proprietary multimodal stack: it introduced MAI-Voice-1 in 2025, then expanded its in-house lineup with MAI-Transcribe-1, MAI-Voice-1 and MAI-Image-2 in April 2026 under an “AI self-sufficiency” strategy. MAI-Transcribe-2 is a rapid iteration in the speech-recognition portion of that effort.
Microsoft had already tied transcription to collaboration products through AI-generated Teams Premium meeting notes. Offering a standalone transcription model through Foundry at a sharply lower early-bird rate extends that capability from a bundled feature toward metered infrastructure for developers.
First-order effects
- Developers using Microsoft Foundry can buy MAI-Transcribe-2 at $0.10 per audio hour through 2026, lowering the direct inference bill for transcription workloads.
- Microsoft AI is using its claimed accuracy advantage over Gemini 3.5 Transcribe and GPT-Transcribe alongside price to compete for speech-recognition deployments.
Second-order effects
- Gemini and GPT transcription offerings face a clearer price-performance comparison, putting pressure on providers to defend premium pricing with measurable accuracy, latency, or workflow advantages.
- Teams-oriented customers and other application builders gain an incentive to separate transcription usage from per-seat collaboration pricing and evaluate it as a usage-based service.
Third-order effects
- If rapid in-house model iterations continue, Microsoft can reduce its dependence on third-party models for speech features while using Foundry distribution to turn model efficiency into a competitive input cost.
- Speech AI is moving toward a market in which accuracy, speed, and inference cost are jointly priced, making model-operating efficiency as important as the feature layer built above it.
The trend: Speech recognition is becoming a metered AI utility, with proprietary model stacks competing on the combined cost, latency, and quality of each transcribed hour.