Mistral launches Small 3, a latency-optimized 24B-parameter model that it says is competitive with larger models such as Llama 3.3 70B or Qwen 32B
Apache 2.0, 81% MMLU, 150 tokens/s — Today we're introducing Mistral Small 3, a latency-optimized 24B-parameter model released under the Apache 2.0 license.
Mistral AI
Related Coverage
- Mistral Small 3 — More notably, they claim the following: Simon Willison's Weblog
- Model Card for Mistral-Small-24B-Base-2501 Hugging Face
- Scaling the Tülu 3 post-training recipes to surpass the performance of DeepSeek V3 Ai2
- 🚀 Announcing Mistral AI Small 3, our most efficient and versatile model yet! — ✅ 24B parameters — ✅ 81% MMLU … Sophia Yang, Ph.D.
- Mistral Small 3 Hacker News
Discussion
-
@lmstudio
@lmstudio
on x
.@MistralAI's Small 3 in now available in LM Studio! 🎉 - 24B params - Apache 2.0 - Over 81% accuracy on MMLU From the terminal, run: “lms get mistral-small-24b” [image]
-
@paul_jacob_
Paul Jacob
on x
Mistral Small 3: - 24B params - 81% MMLU - 150 tokens/s - Competitive with Llama-3.3 70B, Qwen-2.5 32B, GPT4o-mini - Apache 2.0
-
@ollama
@ollama
on x
ollama run mistral-small:24b Mistral Small 3 is here under Apache 2.0!
-
@tres_sarcastik
Léo
on x
@MistralAI This gets me excited. [image]
-
@dchaplot
Devendra Chaplot
on x
Announcing Mistral Small 3! - 24B params, 81% MMLU - Latency optimized: 150 tokens/s - Competitive with Llama-3.3 70B, Qwen-2.5 32B, GPT4o-mini - Apache 2.0 🔗Blog: https://mistral.ai/... 1/N [image]
-
@usb_type_d
@usb_type_d
on x
@MistralAI “No synthetic data” mild shots at deepseek
-
@rseroter
Richard Seroter
on x
You've had your fill of DeepSeek news and are already looking for other AI models to gaze at? @MistralAI Small 3 is out today and offers impressive perf: https://mistral.ai/... @allen_ai just shipped Tülu 3 405B with some terrific perf: https://allenai.org/...
-
@kaggle
@kaggle
on x
🤖 @MistralAI's Small 3 is now available on #KaggleModels! Learn more: https://www.kaggle.com/...
-
@maxwinebach
Max Weinbach
on x
Notice how small models are slowly getting less small? Small was 3-7B parameters, then 14B, now 24B
-
@maxwinebach
Max Weinbach
on x
This goes with my thought that while edge compute is great, these small models won't do it. You'll need 3B specialized models trained to sort/format data in an efficient way to send to a larger model in a cloud. Mixture of Agents.
-
@firstadopter
Tae Kim
on x
I'll say again. If Mistral launched this same model (and they literally did similar efficiency gains with MoE in prior models), no one would have cared because it was France not China.
-
@binarybits
Timothy B. Lee
on x
I don't think “no one would have cared.” It's a genuinely impressive model. But there are exponentially more dumb takes about it (both inflating its significance and calling it a fraud or stolen) than we would have gotten about a mistral model.
-
@mistralai
@mistralai
on x
Introducing Small 3, our most efficient and versatile model yet! Pre-trained and instructed version, Apache 2.0, 24B, 81% MMLU, 150 tok/s. No synthetic data so great base for anything reasoning - happy building! https://mistral.ai/...
-
r/LocalLLaMA
r
on reddit
Mistral Small