Meta launches Muse Voice Transcribe, MSL's first real-time audio perception model, with streaming automatic speech recognition, trained with 70+ languages
Experience Muse Voice Transcribe in real time … We're excited to introduce Muse Voice Transcribe, the first real …
Adding real-time speech recognition gives MSL an audio-input component alongside those media-generation efforts. Alexandr Wang said the model was already powering dictation in the Meta desktop app and Muse Code and was available through Meta's model API.
First-order effects
Meta gains a streaming transcription model trained on more than 70 languages, extending the Muse lineup from query, image, and planned video capabilities into real-time audio perception.
Meta desktop-app and Muse Code users receive the model through dictation, while developers can access it through Meta's model API, according to Wang.
Second-order effects
Meta teams building AI features have a native live-audio input layer rather than relying only on text prompts or separately assembled speech components; Wang described speaker diarization and endpointing as part of the same model.
API availability puts Muse Voice Transcribe in reach of developers whose workflows depend on transcription, increasing the practical value of Meta's broader Muse model portfolio beyond its own consumer surfaces.
Third-order effects
If Meta continues deploying individual Muse components across its apps and API, the competitive unit shifts from a standalone model to a vertically integrated multimodal stack spanning input, generation, and product distribution.
The pattern points toward workflow-native AI in which speech becomes structured application input, not merely an accessibility or dictation feature.
The trend: Meta is building Muse into a multimodal model family that connects real-time inputs and media generation to its applications and developer API.
Meta has released Muse Voice Transcribe, taking the #1 spot for Final Transcript accuracy on AA-WER Streaming with 3.1% WER at 0.16s after end of speech Muse Voice Transcribe is the first streaming Speech to Text model developed by Meta Superintelligence Labs. Meta states that th…
whatever your perspectives on Meta's AI capabilities are, you have to concede there's a new pipeline in place that's churning out interesting work at an impressive clip.
1/ today we're rolling out muse voice transcribe, our first real-time audio perception model - SOTA in streaming speech-to-text. also handles speaker diarization and endpointing natively in a single model.
Understating the strength of that tool for where Meta has an aim in succeeding: computer use The easiest way to get reasoning for desktop or even agentic tasks while crafting from human people is just to have them voice out loud So you need top notch -likely multilingual- audio t…
2/ it's accurate and quick - it uses adaptive delay to decide how much context it needs before transcribing each word. trained on 70+ languages (with 25 validated at launch), handles code-switching mid-sentence, and manages hour+ sessions and 20+ speakers without post-processing.
Muse Voice transcribe is stupid fast and actually good. It nails my Franglais (English/French) perfectly, which nothing else does. I was locked into @superwhisper for my workflow. will try Muse Voice Transcribe. Gj @alexandr_wang
It holds up on messy, real audio too — trained across 70+ languages (with 25 validated at launch), handles mid-sentence code-switching, and manages hour-long sessions with 20+ speakers. You can bias it toward names and domain terms as well.
Already powering dictation in the Meta desktop app and voice input in Muse Code. Live now on the Meta model API with a zero-data-retention tier. https://research.meta.ai/...
🚨 Meta's new voice model looks ridiculously good It's already #1 on Artificial Analysis for speech to text, it only costs $0.18/hour too so it's dirt cheap Meta are grinding right now
I think this is a much bigger Meta AI launch than it looks at first glance. Real-time transcription is already useful, but combining speech-to-text, speaker identification, and endpoint detection into a single model is where things get interesting. That removes a lot of complex…
The model decides when to listen. It waits a little longer on hard words and commits faster on easy ones, using adaptive delay to predict each token and increase accuracy. With adaptive delay, the model is near the pareto front on speed-accuracy tradeoff.
Muse Voice Transcribe is MSL's first real-time audio perception model — rolling out today. SOTA in streaming speech-to-text, it handles speaker diarization, and endpointing natively in a single model.
Super excited about this release! Especially love the fast yet accurate handling of long context, 20+ speakers and mid-sentence code-switching without post-processing!