Nvidia launches Nemotron 3, a family of AI models using a hybrid mixture-of-experts architecture and the Mamba-Transformer design, in 30B, 100B, and ~500B sizes
Nvidia launched the new version of its frontier models, Nemotron 3, by leaning in on a model architecture that the world's …
Subsequent coverage shows Nvidia continuing to build out the line with the 120B-parameter Nemotron 3 Super and a multimodal Nano Omni variant, making this launch the base of a broader Nemotron 3 product family.
First-order effects
Nvidia gains a tiered Nemotron 3 lineup spanning roughly 30B to 500B parameters, giving model users distinct scale options within one architecture family.
The launch puts hybrid mixture-of-experts and Mamba-Transformer design at the center of Nvidia’s own model offering, rather than treating model development solely as a downstream use of its compute platform.
Second-order effects
Competing model providers face added pressure to differentiate on architecture, model scale, and deployment fit as Nvidia expands a branded family rather than a single release.
Organizations evaluating Nvidia’s AI stack can compare successive Nemotron 3 variants—including the later multimodal Nano Omni model—without changing model families, potentially concentrating experimentation around Nvidia’s ecosystem.
Third-order effects
If the family continues to expand, frontier-model competition may increasingly center on heterogeneous architectures tailored to different capability and efficiency trade-offs, not parameter count alone.
Nvidia’s model releases point toward a more integrated AI stack in which a compute supplier also shapes the model layer; the degree of customer lock-in will depend on how portable and competitive those models remain.
The trend: This is part of the shift from monolithic frontier models toward segmented, hybrid model families designed for different AI workloads and deployment constraints.
NVIDIA has just released Nemotron 3 Nano, a ~30B MoE model that scores 52 on the Artificial Analysis Intelligence Index with just ~3B active parameters Hybrid Mamba-Transformer architecture: Nemotron 3 Nano combines the hybrid Mamba-Transformer approach @NVIDIAAI has used on [ima…
NVIDIA focused on efficiency as well as intelligence with Nemotron 3 Nano, and it presents an attractive trade-off between speed and capability. In pre-release testing of the @DeepInfra serverless endpoint, we saw output speeds of ~380 tokens per second [image]
✨ Meet our new open family of models: @NVIDIA Nemotron 3 Open in weights, data, tools, and training, Nemotron 3 is built for multi-agent apps and features: • An efficient hybrid Mamba‑Transformer MoE architecture • 1M token context for long-term memory and improved reasoning [vid…
It's an honor to be competing with Nvidia for the best models with open data, checkpoints, and code. Super excited about Nemotron 3 and Nvidia's new focus on fully open models in 2025.
New: Nemotron v3 is open, fastest, highest benchmark scoring. Nemotron v3 Nano delivers 4x higher throughput than Nemotron 2 Nano & delivers most tokens per second at scale using hybrid mamba/transformer MoE architecture - state space models are the way! https://research.nvidia.c…
Today, @NVIDIA is launching the open Nemotron 3 model family, starting with Nano (30B-3A), which pushes the frontier of accuracy and inference efficiency with a novel hybrid SSM Mixture of Experts architecture. Super and Ultra are coming in the next few months. [image]