Meta unveils open-source AI models the company says can identify 4,000+ languages and produce speech for 1,000+ languages, a 40x and 10x increase, respectively
They could help lead to speech apps for many more languages than exist now. — Meta has built AI models that can recognize …
MIT Technology ReviewRhiannon Williams
Context & Ripple Effects
Meta had already open-sourced a translation model covering 200 languages as part of its universal speech-translator effort, a baseline this release substantially expands through its earlier 200-language translation model.
The announcement shifts that effort from translation coverage toward broader speech recognition and synthesis. It also precedes Meta's later SeamlessM4T multimodal translation and transcription release, suggesting a continuing buildout of shared multilingual speech infrastructure.
First-order effects
Developers and researchers can access Meta's models for recognizing more than 4,000 languages and generating speech in more than 1,000, lowering the model-access barrier for language-specific speech applications.
Meta strengthens its position in multilingual AI research by publishing capabilities that extend well beyond its earlier 200-language open-source translation model.
Second-order effects
Speech-app builders can test support for languages that may have lacked practical recognition or voice-generation tooling, while needing to validate quality for each language and use case.
Rival model providers face added pressure to compete on language coverage and openness; Meta's subsequent SeamlessM4T release indicates that coverage can become a platform for translation and transcription products.
Third-order effects
If broad language coverage is paired with usable quality and open access, multilingual speech technology may increasingly be built on a small set of shared foundation models rather than separate language-by-language systems.
The competitive boundary could move from raw language coverage toward distribution, product integration, and the data and evaluation needed to serve less-represented languages reliably.
The trend: This is part of the push to make multilingual speech AI a broadly available foundation layer, expanding language coverage through open models rather than limiting it to major-language products.
All you need to build the Tower of Babel is a single model that supports 1000s of spoken languages. And use the Bible for training, literally. Meta hits another remarkable Llama milestone for speech...
MMS: Massively Multilingual Speech. - Can do speech2text and text speech in 1100 languages. - Can recognize 4000 spoken languages. - Code and models available under the CC-BY-NC 4.0 license. - half the word error rate of Whisper. Code+Models: https://github.com/... Paper:... http…
😭 This is beautiful, it even supports text-to-speech and asr in tarifit, in both arabic and latin scripts... Clearly this was built with love 😭 https://twitter.com/... [image]
Nothing reaps better returns than Attention + Ads with AI at scale! Search is #2, commerce/recsys #3 | still think Meta is behind ? And nothing hits scale better than open source. Open source models have advantages around community-driven innovation, cost management, and trust...…
🇮🇱🇯🇵🇩🇪🇪🇸🇮🇳 x 1000 This is MASSIVE folks! (blind reaction) TTS and STT in one model, that understands 1100 languages, better than whisper! and is able to generate audio in those languages? Incredible thanks to @ylecun @boztank and tons of other folks who made this happen and relea…
New work! The Massively Multilingual Speech (MMS) project scales speech technology to 1,100-4,000 languages using self-supervised learning with wav2vec 2.0. Paper: https://research.facebook.com/ ... Blog: https://ai.facebook.com/... Code/models: https://github.com/... [video]
So smart, Meta's new massively multilingual speech model was trained on New Testament biblical readings! This is a really interesting angle as religions are the stories that survive and proliferate.. it's an interesting form of grounding and can be expanded upon https://twitter.c…
Meta is the only one using permissive licensing for their models out of the big companies. 🤖Llama gave us Alpaca and then OpenLLaMA They are going all in, and they are nailing it. Open source is the only way. https://twitter.com/...
This is big news. Among its benefits, MMS can empower marginalized communities by overcoming linguistic barriers and preserving endangered languages. #ArtificialIntelligence #SpeechRecognition #Inclusion https://twitter.com/...
@ylecun Oh wow this is huge ! This model covers dialects for which it was impossible to build a strong dataset ! But somehow with Meta's huge conversational base, it became possible ! Great work Meta ! And great work @ylecun 🙌
Another announcement coming out of @MetaAI 🚨 Mark Zuckerberg just announced that they are open sourcing and introducing ‘Massively Multilingual Speech’. We could first only identify up to 100 languages online via software. Now? 4000. https://twitter.com/... [video]
more open source AI efforts from meta,,, this time the massively multilingual speech project, an attempt to use machine learning to provide speech to text (and vice versa) to the thousands of languages spoken (or no longer spoken) in the world https://ai.facebook.com/...