Mistral AI launches Mixtral 8x22B, its latest sparse mixture-of-experts model, after releasing Mixtral 8x7B in December 2023
As Google unleashed a barrage of artificial intelligence announcements at its Cloud Next conference, Mistral AI decided to jump into action with the launch …
VentureBeatShubham Sharma
Context & Ripple Effects
Mixtral 8x22B advances Mistral’s sparse mixture-of-experts line after the earlier 8x7B release, giving the company a new model-family milestone amid a period of prominent AI platform announcements.
Mistral adds Mixtral 8x22B to its sparse mixture-of-experts offerings, expanding the options associated with the Mixtral family.
Developers and organizations evaluating Mistral models gain another named model configuration to assess alongside Mixtral 8x7B.
Second-order effects
The launch puts additional pressure on rival model providers to differentiate not only on model scale but also on architecture and workload fit.
As Mistral broadens its catalog, buyers face a more segmented evaluation process—comparing general-purpose, specialized, and later multimodal options rather than selecting a single vendor model.
Third-order effects
If this portfolio pattern persists, foundation-model competition will increasingly center on a mix of architectures and task-specific families, not a single flagship benchmark race.
That shift could favor providers that can package models into coherent deployment choices; the corpus does not establish which architecture will become the default.
The trend: Mixtral 8x22B is an early point in the trend toward diversified AI model portfolios organized by architecture, capability, and deployment fit.
And seems like we have a first conversion to transformers of the new 146B (8x22b) parameters models Mistral released a few hours ago: https://huggingface.co/... Congrats LagPixelLOL (v2ray)! For more community version, search for 8x22b on the hub: https://huggingface.co/... How m…
Mixtral 8x22B running on a MacBook Pro with Ollama Works with the latest pre-release version of 0.1.32 and will be published to https://ollama.com/... soon. [video]
sooo DBRX got to be the best open-weights model for exactly two weeks 😭 and the new mistral model will probably last another few weeks til Llama-3 (fingers crossed!) 🤐 progress in open models is crazy!
Apparently the new Mistral model beats Claude Sonnet and is a tad bit worse than GPT-4 In a couple of months, the open source community will fine tune it to beat GPT-4 This is a fully open weights model with an Apache 2 license! I can't believe how quickly the OSS community...
Mistral's surprise model is 8x22B which is 176billion params, like GPT3.5 Very excited about this given that Mixtral, at 56B, is already close/surpassing 3.5 So new mistral will be even better, perhaps approaching gpt4
New Mixtral 8x22B runs nicely in MLX on an M2 Ultra. 4-bit quantized model in the 🤗 MLX Community: https://huggingface.co/... h/t @Prince_Canuma for MLX version and v2ray for HF version https://huggingface.co/v2ray [video]
New Mixtral 8x22b released on Google Cloud Next day. With a reasonable quantization one can safely run them with 4x{A,H}100 cards. Cannot wait to see its detailed quality in comparison with other state of the art models. (Yes, strictly speaking you can run with 3 cards, but the..…
mixtral 8x22B - things we know so far 🫡 > 176B parameters > performance in between gpt4 and claude sonnet (according to their discord) > same/ similar tokeniser used as mistral 7b > 65536 sequence length > 8 experts, 2 experts per token: More > would require ~260GB VRAM in... [im…
Can't download @MistralAI's new 8x22B MoE, but managed to check some files! 1. Tokenizer identical to Mistral 7b 2. Mixtral (4096,14336) New (6144,16K), so larger base model used. 3. 16bit needs 258GB of VRAM. BnB 4bit 73GB. HQQ 4bit attention, 2 bit MLP 58GB VRAM => H100 fits!..…