/
Navigation
Chronicles
Browse all articles
Explore
Semantic exploration
Research
Entity momentum
Nexus
Correlations & relationships
Story Arc
Topic evolution
Drift Map
Semantic trajectory animation
Posts
Analysis & commentary
Pulse API
Tech news intelligence API
Browse
Entities
Companies, people, products, technologies
Domains
Browse by publication source
Handles
Browse by social media handle
Detection
Concept Search
Semantic similarity search
High Impact Stories
Top coverage by position
Sentiment Analysis
Positive/negative coverage
Anomaly Detection
Unusual coverage patterns
Analysis
Rivalry Report
Compare two entities head-to-head
Semantic Pivots
Narrative discontinuities
Crisis Response
Event recovery patterns
Connected
Search: /
Command: ⌘K
Embeddings: large
TEXXR

Chronicles

The story behind the story

days · browse · Enter similar · o open

Meta open sources its AI translation model, which works across 200 languages, as part of its ambitious R&D project to create a “universal speech translator”

helping editors to make more information available in under-represented languages. (1/2) https://ai.facebook.com/... @wiki_early_life : If AI Fairness researchers were serious they would spend more time celebrating this (which makes huge amounts of text more accessible to hundreds of millions of people) than complaining about popcorn farts worth of CO2 emissions to train GPT-3 https://twitter.com/... Andrew Tschesnok / @tschesnok : @Suhail But is it not exclusively in the domain of the well funded? I know of startups raising $40m primarlay for compute costs. @metaai : (4/4) 🎉NLLB has the potential to benefit teams and technologies across Meta, and as we continue to build upon it, we look forward to doing just that. Take a deeper dive into the project here: https://ai.facebook.com/... https://twitter.com/... @metaai : 🆕🆕🆕❗ ️NLLB-200 Model, Evaluation Dataset, and Paper, with improved translation quality for over 200 languages. Help advance AI translation for low-resource languages. 📃Paper: https://research.facebook.com/ ... ⚙️Model: https://github.com/... 🛤️Benchmark: https://github.com/... https://twitter.com/... Olivier / @olli757 : @tschesnok @Suhail the actual models they released today are available for download. No need to train them yourself and spend millions on computing resources. @metaai : @facebook ... Discover how our No Language Left Behind project is driving inclusion through the power of AI translation: https://research.facebook.com/ ... John Koetsier / @johnkoetsier : Mark Zuckerberg wants to build a Star Trek-like “universal translator” >> Meta open sources early-stage AI translation tool that works across 200 languages https://www.theverge.com/... Stephen Shankland / @stshank : This is a pretty mammoth amount of translation technology that Facebook has open-sourced. https://t.co/jBz0fdw2Wn Boz / @boztank : Huge kudos to our AI teams behind No Language Left Behind, a single AI model that can now translate 200 different languages, many of which were previously unsupported by translation systems. https://www.youtube.com/... Yann LeCun / @ylecun : “No Language Left Behind.” An open-source language translation system from FAIR capable of translating 200 languages between each other. 50 billion parameters. The code and models are made available today as part of the Fairseq package. Github: https://t.co/1m581UAfpo 1/n https://t.co/NCfTzyJWqA

The Verge James Vincent

Context & Ripple Effects

Meta is making the NLLB-200 model, its evaluation dataset, and supporting paper available rather than keeping the work solely inside its research program. That gives the project a public research and development base focused on language coverage that is often underrepresented in translation tools.

The release became an early step in a broader Meta language-AI arc: later coverage moved into combined text-and-speech translation and transcription and a much larger open-source language and speech effort.

First-order effects

  • Researchers and developers can access Meta's translation model and evaluation materials for work spanning 200 languages, instead of treating the project as a closed internal capability.
  • Editors seeking to publish information in underrepresented languages gain a readily available model base for translation workflows.

Second-order effects

  • The public evaluation dataset gives language-AI developers a shared basis for comparing coverage and quality across languages, increasing pressure to demonstrate performance beyond the most commonly supported ones.
  • Meta's release establishes reusable infrastructure for its own subsequent move from text translation toward translation and transcription across text and speech.

Third-order effects

  • If this release pattern persists, multilingual AI competition shifts from isolated translation products toward open model-and-dataset stacks that support speech as well as text.
  • Broader language coverage makes model access a more consequential part of international AI deployment, because the distribution of capable models influences which language communities can build on them.

The trend: Multilingual AI is evolving from closed, text-centric translation systems into openly distributed model stacks designed to cover more languages and modalities.

Discussion

  • @suhail @suhail on x
    @Olli757 @tschesnok Exactly - people will pretrain from this for other weird use cases now unrelated to translation. AI building blocks.
  • @an_open_mind Jerome Pesenti on x
    Awesome new work by @MetaAI that allows translating 200 languages in a single open sourced model - NLLB-200 - available to all and already used by Wikipedia Page: https://ai.facebook.com/... Blog: https://ai.facebook.com/... Paper: https://research.facebook.com/ ... Github: https…
  • @bensprecher Ben Sprecher on x
    Wow. Some serious work here. Bravo.... Looking forward to reading the paper in detail.... https://twitter.com/...
  • @yudhanjaya Yudhanjaya Wijeratne on x
    For years we've been pointing out that social media sites, as massive text repositories, should lead the charge on language translation. Really glad to see more open source work coming out from Meta on the subject. This is fantastic. https://twitter.com/...
  • @metaai @metaai on x
    (3/4) 💻@Wikipedia editors are using our technology to translate articles in 20+ low-resource languages, including 10 that previously were not supported by any machine translation tools on their platform like Luganda. https://ai.facebook.com/...
  • @wikimedia @wikimedia on x
    Wikipedia editors have been piloting @MetaAI's model to translate articles in 20+ low-resource languages (those without extensive datasets to train AI systems), including 10 that were not supported by any machine translation tools before. (2/2) https://ai.facebook.com/...
  • @minderellasf Miranda Duncan on x
    FYI. I'm kind of excited about this. I use automated translation technologies rather regularly as an expat yogi in Asia, sometimes to order things , like via transport service app, & with usually humor when top news stories endeavoring to decipher at least a shred of meaning. 💙 h…
  • @suhail @suhail on x
    @tschesnok The biggest insight these past few years is that models are starting to increasingly commoditize. Go to Hugging Face and you'll see tons to pre-trained ones to get started.
  • @boardsofdata Martin Palazzo on x
    50B parameters 👁️👄👁 ️. Dall-e tiene 12B, Bert 300 million. https://twitter.com/...
  • @suhail @suhail on x
    If you were looking for a revolution, AI is early days with continuous breakthroughs. Article: https://research.facebook.com/ ... https://twitter.com/...
  • @wikimedia @wikimedia on x
    Since 2014, @Wikipedia editors have used our Content Translation Tool to share knowledge in more languages. Now, the tool includes open-source technology from @MetaAI—helping editors to make more information available in under-represented languages. (1/2) https://ai.facebook.com/…
  • @wiki_early_life @wiki_early_life on x
    If AI Fairness researchers were serious they would spend more time celebrating this (which makes huge amounts of text more accessible to hundreds of millions of people) than complaining about popcorn farts worth of CO2 emissions to train GPT-3 https://twitter.com/...
  • @tschesnok Andrew Tschesnok on x
    @Suhail But is it not exclusively in the domain of the well funded? I know of startups raising $40m primarlay for compute costs.
  • @metaai @metaai on x
    (4/4) 🎉NLLB has the potential to benefit teams and technologies across Meta, and as we continue to build upon it, we look forward to doing just that. Take a deeper dive into the project here: https://ai.facebook.com/... https://twitter.com/...
  • @metaai @metaai on x
    🆕🆕🆕❗ ️NLLB-200 Model, Evaluation Dataset, and Paper, with improved translation quality for over 200 languages. Help advance AI translation for low-resource languages. 📃Paper: https://research.facebook.com/ ... ⚙️Model: https://github.com/... 🛤️Benchmark: https://github.com/... ht…
  • @olli757 Olivier on x
    @tschesnok @Suhail the actual models they released today are available for download. No need to train them yourself and spend millions on computing resources.
  • @metaai @metaai on x
    @facebook ... Discover how our No Language Left Behind project is driving inclusion through the power of AI translation: https://research.facebook.com/ ...
  • @johnkoetsier John Koetsier on x
    Mark Zuckerberg wants to build a Star Trek-like “universal translator” >> Meta open sources early-stage AI translation tool that works across 200 languages https://www.theverge.com/...
  • @stshank Stephen Shankland on x
    This is a pretty mammoth amount of translation technology that Facebook has open-sourced. https://t.co/jBz0fdw2Wn
  • @boztank Boz on x
    Huge kudos to our AI teams behind No Language Left Behind, a single AI model that can now translate 200 different languages, many of which were previously unsupported by translation systems. https://www.youtube.com/...
  • @ylecun Yann LeCun on x
    “No Language Left Behind.” An open-source language translation system from FAIR capable of translating 200 languages between each other. 50 billion parameters. The code and models are made available today as part of the Fairseq package. Github: https://t.co/1m581UAfpo 1/n https:/…