Mistral announces Mistral Large 2, the new generation of its flagship model, with 123B parameters; commercial usage requires a separate license
Today, we are announcing Mistral Large 2, the new generation of our flagship model. Tobias Mann / The Register : Mistral Large 2 leaps out as a leaner, meaner rival to GPT-4-class AI models MD Ijaj Khan / Hindustan Times : Google integrates Mistral AI's codestral model: What is it and how will it help developers? The Indian Express : Mistral unveils new AI model Large 2, takes on Meta's Llama 3.1 and OpenAI's GPT-4o Luke Jones / WinBuzzer : Mistral Launches 123 Billion Parameter “Mistral Large 2” AI Model Ben's Bites : Mistral adds Large 2 to frontier AI models. Alexey Shabanov / TestingCatalog : Mistral Large 2 now available on major cloud platforms and Le Chat Pradeep Viswanathan / Neowin : Mistral announces Large 2 flagship LLM with 123 billion parameters Threads: @altstoreio : We've explicitly reviewed these initial sources to ensure they meet our safety standards. Just add them from the Sources screen and new apps will appear in your Browse tab! We've got a great first batch for y'all, keep reading to learn more 👇 X: Devendra Chaplot / @dchaplot : Super excited to announce Mistral Large 2 - 123B params - fits on a single H100 node - Natively Multilingual - Strong code & reasoning - SOTA function calling - Open-weights for non-commercial usage Blog: https://mistral.ai/... Weights: https://huggingface.co/... 1/N [image] Ravid Shwartz Ziv / @ziv_ravid : Mistral released a new 123B dense model, which is very good. Why is it dense and not a mixture of experts? There was a time a few months ago when it seemed that a mixture of experts was the promising direction, but it doesn't look like that anymore; why? LangChain / @langchainai : 🧠Mistral Large 2: improved reasoning and function calling 🧠 Try out the next generation of the @MistralAI flagship model, with its significantly improved reasoning, function calling, and more. Mistral announcement: https://mistral.ai/... LangChain PY docs: [image] Simon Willison / @simonw : Anyone know what benchmark @MistralAI are reporting on here for “Function Calling” for their new Mistral Large 2 model? Their post doesn't name the benchmark used: https://mistral.ai/... [image] @rajko_rad : What a week!!! it's amazing to have 405B open model but realistically you can't deploy that with in prod - here's the new SoTa deployable model 😎🔥 [image] Rowan Cheung / @rowancheung : Wow. After only one day of Llama 3.1 405b, French startup Mistral AI dropped LARGE 2. It's ANOTHER open-source flagship AI model that scores close to Llama 3.1 405b and even surpasses it on coding benchmarks while being much smaller at 123b. Benchmarks vs. Llama 3.1 405b: - [image] Richard Kelley / @richardkelley : Interesting to see “fits on a single GPU” as a selling point: @bilaltwovec : props to mistral for keeping w the efficiency corner branding after the last time [image] @teknium1 : btw now we have a 4th model in the public sphere comparable to gpt4. It's also like a spectrum of full open access to fully closed now lol Llama-3.1 - Fully Open (yea not fully open source) Mistral Large 2 - Open but non-commercial Sonnet 3.5 - closed GPT4 - closed [image] @teknium1 : Mistral just released a 123B param instruct tune only model called Mistral Large, unfortunately is a non-commercial license, and no base model was released, but looks great for personal use: Model: https://huggingface.co/... Blog Post: https://mistral.ai/... @guillaumelample : Today, we release Mistral Large 2, the new version of our largest model. Mistral Large 2 is a 123B-parameter model with a 128k context window. On many benchmarks (notably in code generation and math), it is superior or on par with Llama 3.1 405B. Like Mistral NeMo, it was trained @mistralai : https://mistral.ai/...
Context & Ripple Effects
Mistral had already positioned its first Large model as a lower-cost GPT-4 alternative with a 32K context window; Large 2 extends that flagship line with a substantially longer context capacity and a different access boundary for business users.
The release follows a split product strategy: Mistral and Nvidia had made the 12B Mistral NeMo available under Apache 2.0, while Codestral's commercial-use restrictions had already shown that Mistral would not apply one licensing model across its portfolio.
First-order effects
- Mistral gives developers and cloud customers a 123B-parameter flagship that it says can run on a single H100 node, while making commercial deployment contingent on a separate license.
- Non-commercial users can access the released weights, but companies must evaluate licensing alongside Large 2's model performance, 128K context window, and function-calling capabilities.
Second-order effects
- Large 2 sharpens the choice for enterprise buyers between proprietary API access, open-weight experimentation, and separately licensed deployment—especially where infrastructure control matters.
- Mistral's cloud availability makes distribution partners immediate routes to market, while its commercial terms preserve a direct negotiating lever rather than treating released weights as a fully open commercial product.
Third-order effects
- If this mixed-access approach persists, frontier-model competition may increasingly separate weight availability from commercial rights: developers gain evaluation and customization options, while vendors retain monetization control over production use.
- Efficiency claims such as fitting a 123B model on one H100 node raise the strategic value of deployable models, not only benchmark-leading scale; whether that changes buyer behavior depends on real-world cost and performance.
The trend: Frontier AI vendors are pairing more deployable models with segmented licensing to compete for enterprise adoption without fully surrendering control of commercial use.