Meta AI unveils an open-source translator for spoken but unwritten languages, which currently translates one sentence at a time between Hokkien and English
Meta AI built the first speech translator that works for languages that are primarily spoken rather than written. Andrew Hutchinson / Social Media Today : Meta Develops New Speech-to-Speech Real-Time Audio Translation Process Kevin Hurler / Gizmodo : Meta Has Developed AI for Real-Time Translation of Hokkien Kyt Dotson / SiliconANGLE : Meta builds AI-powered speech translation for Hokkien to understand unwritten languages Tech Xplore : Meta touts AI that translates spoken-only language Tweets: George Hadjia / @ghadjia : This is pretty sick. $META https://twitter.com/... Boz / @boztank : One step closer to the universal translator! Thousands of languages around the world have no standard written form, and today our AI team released a breakthrough in translating them: the first Hokkien language translation system https://ai.facebook.com/... @quinnypig : Meta's apparently focusing on languages that are spoken and not written, like “internal Facebook discussions that even slightly touch on antitrust issues.” https://twitter.com/... Ti-Chung Cheng / @tichungcheng : As much as Meta has their issues with the world, and challenges that HAI still needs to address, I applaud their contribution at researching ways to translate dialects like Hokkien (Taiwanese). I would never want to lose one of my mother tongues. https://twitter.com/... Jane Manchun Wong / @wongmjane : This is great — if @MetaAI's UST also supports Cantonese in the future. I'll be able to communicate with my relatives in my native Cantonese while they speak in their native Hokkien and that we understand each other https://twitter.com/... Emily Y. Wu / @emilyywu : Hokkien (where Taiwanese Hokkein / Taigi developed from) is the demo lang for Facebook's AI speech translation project. The video of Mark Zuckerberg conversing with his Hokkein-speaking engineer tho, is on FB only: https://m.facebook.com/... https://twitter.com/... Joshua Yang / @joshiunn : Huge. Meta researchers developed a speech-to-speech translator for #Taiwanese Hokkien, using 30k hrs of TW drama data. Maybe there's a metaverse where mini Zuckerbergs just watch Taiwan drama w/ my a-má, all discussing how evil that one character truly is. https://huggingface.co/... https://twitter.com/... https://twitter.com/... Lawrence Lundy-Bryan / @lawrencelundy : Common now we did it guys, we finally got a bablefish. Now how power hungry are these models and how quickly can we get them in earpods? https://twitter.com/... Scott Redler / @reddogt3 : The future when Markets are ready https://twitter.com/... Clark / @canteringclark : Imagine a future where you get some kind of small cochlear implant that translates all languages. Think of the ways this could improve the world. You don't realize how much language divides common people. https://twitter.com/... @msjoycelee : @GHadjia Not gonna lie, this is amazing. Just want to clarify something that Mark said though: Hokkien/Taiwanese is a dialect and that we are able to read traditional or simplified Chinese and say it in Hokkien/Taiwanese. So, the “standard writing system” is just written Chinese.
Context & Ripple Effects
Meta had already open-sourced a 200-language translation model as part of its universal speech-translator effort. The Hokkien-to-English release extends that work to a language treated here as primarily spoken rather than written, shifting the technical focus from text coverage to speech data and direct speech translation.
Later Meta releases—SeamlessM4T’s text-and-speech translation and transcription and the Seamless Communication suite—place this narrow Hokkien system at the beginning of a broader effort to make multilingual communication more natural across modalities.
First-order effects
- Meta makes a Hokkien-to-English speech-translation system available as open source, giving researchers and developers a starting point for work on spoken-first languages.
- Hokkien and English speakers gain a demonstrable translation path, although the reported system operates one sentence at a time rather than as continuous conversation.
Second-order effects
- Translation-model developers are pushed to compete on speech coverage and training methods for languages without standard written corpora, not only on the number of text languages supported.
- Meta’s open release can make its approach a shared research baseline, while later products such as Seamless Communication can build differentiation through broader capability and more natural interaction.
Third-order effects
- If speech-first translation methods generalize, language AI will be evaluated less by text-language counts and more by whether it serves communities excluded by text-centric datasets.
- Open model releases may separate foundational translation research from the consumer interfaces that deploy it, as Meta’s later plan to bring Meta AI into its apps illustrates.
The trend: Translation AI is moving from broad text-language coverage toward speech-native systems designed to include languages with limited written data.