Nvidia and Mistral release Mistral NeMo, a 12B-parameter language model with a 128K-token context window, available under the Apache 2.0 open-source license
Mistral NeMo: our new best small model. A state-of-the-art 12B model … Jonathan Kemper / The Decoder : Mistral releases three new LLMs for math, code and general tasks X: Prince Canuma / @prince_canuma : Mistral NeMo Instruct (Q4) is blazing fast 🔥 Running locally on M3 Max at 37 tokens/s using MLX 🚀 > pip install fastmlx And install mlx-lm from source (PR #895) [video] Anshel Sag / @anshelsag : This is a very interesting partnership and one that will likely only increase NVIDIA's relevance. Simon Willison / @simonw : “Mistral NeMo: our new best small model. A state-of-the-art 12B model with 128k context length, built in collaboration with NVIDIA, and released under the Apache 2.0 license.” Sounds like there's a new OpenAI model coming later too, a 4o-mini replacement for GPT-3.5 Turbo Sankalp / @dejavucoder : >128k context window >beats competitors in benchmarks >multi-lingual; >new tiktoken-based tokenizer called tekken trained over multiple languages, >weights on hf [image] @reach_vb : Let's goooo! Nvidia & Mistral release Mistral NeMo 12B 🔥 > Apache 2.0 licensed w/ 128K contact > Beats Llama 3 8B, Gemma 2 9B > Multilingual - EN, FR, DE, ES, IT, PT, CN, JP, KR, AR ⚡️ > Base + Instruct model (uncensored) > Supports function calling > Trained on both [image] Alex Volkov / @altryne : Mistral team is on fire this week damn Colab with @nvidia releasing a 128k!! 12B new open source model that's likely to be the next small-ish OSS darling called NeMo 👏 [image] Arthur Mensch / @arthurmensch : Today, we're announcing Mistral NeMo, a tiny multilingual model, 128k context length, trained with quantization awareness in collaboration with the NVIDIA research team. Kyle Russell / @kylebrussell : “Mistral NeMo was trained with quantisation awareness, enabling FP8 inference without any performance loss.”
Context & Ripple Effects
Mistral had already positioned itself against larger proprietary systems with Mistral Large's lower-cost, 32K-context offering. NeMo extends that arc toward a smaller model tier while materially expanding the amount of text a deployment can keep in context.
The release also differs from Mistral's earlier coding model, whose commercial use was restricted, by making the new weights available under Apache 2.0. That licensing choice makes the Nvidia research collaboration consequential beyond a single hosted product.
First-order effects
- Developers and enterprises can obtain and adapt a 12B model with a 128K-token context window under Apache 2.0, rather than being limited to an API or a restrictive commercial-use license.
- Nvidia gains a directly associated open-weight model whose quantization-aware training supports FP8 inference, tying the model's deployment story to efficient inference hardware and software.
Second-order effects
- Smaller-model vendors face added pressure to pair local-deployment economics with long-context capability; the reported comparisons with 8B- and 9B-parameter rivals make benchmark positioning part of that competition.
- Open availability shifts value toward deployment tooling, fine-tuning, support and inference optimization, since users can take the weights to their own environments.
Third-order effects
- If more capable small models are released permissively, model suppliers may increasingly differentiate through ecosystem distribution, infrastructure performance and specialized services rather than access to weights alone.
- Long-context capability is becoming a practical design variable for smaller models, though its real cost and usefulness will depend on each deployment's inference setup and workload.
The trend: This is one data point in the shift toward permissively licensed, inference-efficient models that broaden buyer choice while moving competition into deployment infrastructure and surrounding services.