Mistral launches Small 3, a latency-optimized 24B-parameter model that it says is competitive with larger models such as Llama 3.3 70B or Qwen 32B
Apache 2.0, 81% MMLU, 150 tokens/s — Today we're introducing Mistral Small 3, a latency-optimized 24B-parameter model released under the Apache 2.0 license.
Mistral AI
Context & Ripple Effects
Mistral had already split its lineup between an Apache-licensed 12B model, Mistral NeMo, and its separately licensed 123B flagship, Mistral Large 2. Small 3 places a substantially larger model on the open-license side while making latency a central product claim.
The release is an early step in Mistral's continuing small-model line: the later Small 3.1 update retained the 24B scale while adding multimodal and multilingual capabilities. That makes Small 3 a meaningful baseline for how the company packages efficient open models.
First-order effects
Developers can use, modify, and distribute the 24B Small 3 under Apache 2.0, expanding self-hosted deployment options without a separate commercial-use license.
Mistral gives buyers a latency-oriented alternative to larger cited models, with reported 81% MMLU and 150-token/s throughput as the basis for evaluation.
Second-order effects
Organizations comparing Llama and Qwen deployments gain another candidate for routing lower-latency workloads, increasing pressure to benchmark performance per serving requirement rather than parameter count alone.
Mistral's own product segmentation becomes clearer: open, efficient models can address broad deployment needs while larger or more specialized offerings remain differentiated by capability and licensing.
Third-order effects
If performance at smaller scales continues to narrow the gap with larger systems, model selection will increasingly hinge on inference latency, operational fit, and licensing—not just benchmark leadership.
Apache-licensed weights can strengthen buyer leverage by making multi-model and self-hosted strategies more practical, though real-world quality and infrastructure costs will determine adoption.
The trend: This is part of the shift toward efficient open-weight models competing for production workloads where inference performance and deployment freedom matter as much as raw scale.
.@MistralAI's Small 3 in now available in LM Studio! 🎉 - 24B params - Apache 2.0 - Over 81% accuracy on MMLU From the terminal, run: “lms get mistral-small-24b” [image]
You've had your fill of DeepSeek news and are already looking for other AI models to gaze at? @MistralAI Small 3 is out today and offers impressive perf: https://mistral.ai/... @allen_ai just shipped Tülu 3 405B with some terrific perf: https://allenai.org/...
This goes with my thought that while edge compute is great, these small models won't do it. You'll need 3B specialized models trained to sort/format data in an efficient way to send to a larger model in a cloud. Mixture of Agents.
I'll say again. If Mistral launched this same model (and they literally did similar efficiency gains with MoE in prior models), no one would have cared because it was France not China.
I don't think “no one would have cared.” It's a genuinely impressive model. But there are exponentially more dumb takes about it (both inflating its significance and calling it a fraud or stolen) than we would have gotten about a mistral model.
Introducing Small 3, our most efficient and versatile model yet! Pre-trained and instructed version, Apache 2.0, 24B, 81% MMLU, 150 tok/s. No synthetic data so great base for anything reasoning - happy building! https://mistral.ai/...