/
Navigation
Chronicles
Browse all articles
Explore
Semantic exploration
Research
Entity momentum
Nexus
Correlations & relationships
Story Arc
Topic evolution
Drift Map
Semantic trajectory animation
Posts
Analysis & commentary
Pulse API
Tech news intelligence API
Browse
Entities
Companies, people, products, technologies
Domains
Browse by publication source
Handles
Browse by social media handle
Detection
Concept Search
Semantic similarity search
High Impact Stories
Top coverage by position
Sentiment Analysis
Positive/negative coverage
Anomaly Detection
Unusual coverage patterns
Analysis
Rivalry Report
Compare two entities head-to-head
Semantic Pivots
Narrative discontinuities
Crisis Response
Event recovery patterns
Connected
Search: /
Command: ⌘K
Embeddings: large
TEXXR

Chronicles

The story behind the story

days · browse · Enter similar · o open

Nvidia releases Nemotron 3.5 Lightning, an open 30B-parameter MoE model, and NeMo Switchyard, an open-source model routing library for AI agents

SiliconANGLE Kyt Dotson

Context & Ripple Effects

Nvidia’s Nemotron line began as a hybrid MoE model family spanning several sizes, then added a 120B open-weight Super model and a multimodal Nano Omni variant. The new release extends that sequence with a 30B model focused on output speed.

The addition of an open-source routing library matters because it pairs Nvidia’s model releases with software for directing AI-agent requests, rather than treating the model as the only product layer.

First-order effects

Second-order effects

  • Model providers and inference vendors serving agent builders face a more integrated Nvidia offer: an open model plus routing software can reduce the need for developers to assemble those layers separately.
  • Nvidia’s routing layer makes inference performance a more direct product-selection criterion for agent developers, aligning with its non-exclusive rights to Groq’s inference technology.

Third-order effects

  • If Nvidia continues coupling open Nemotron models with orchestration tools, competition shifts from standalone model releases toward integrated stacks that control both model access and runtime routing.
  • Open routing software may give agent builders more leverage to switch among models, while concentrating strategic value in the infrastructure that determines where requests run.

The trend: AI model vendors are moving from releasing weights alone toward providing the routing and inference layers that shape how agent workloads choose models.

Discussion

  • @artificialanlys @artificialanlys on x
    NVIDIA has just released the first Nemotron 3.5 model: Nemotron 3.5 Lightning, a highly efficient small open weights model with performance similar to gpt-oss-120b at around a quarter of the total parameters Nemotron 3.5 Lightning is the successor to @nvidia Nemotron 3 Nano 30B […
  • @fastinoai @fastinoai on x
    In collaboration with @nvidia we're releasing two new open weight models: Fastino-Nemotron-3.5-Lightning- Finance and Fastino-Nemotron-3.5-Lightning- Healthcare. Working closely with the Nemotron team, we developed both models on Nemotron 3.5 Lightning using the Fastino [image]
  • @pidotdev @pidotdev on x
    Nvidia Nemotron 3.5 Lightning delivers leading accuracy on agentic coding tasks and up to 4x the speed of comparable open models. With open weights, datasets, and recipes, Nemotron is open and easy to customize. Welcome to Pi, Nemotron 3.5 Lightning! [image]
  • @jensenhuang Jensen Huang on x
    Lightning strikes for continuous and long-run agents! Nemotron 3.5 Lightning is smart, fast, efficient and open.
  • @kimmonismus @kimmonismus on x
    One day after Meta released Muse Glimmer, NVIDIA launched Nemotron 3.5 Lightning and NeMo Switchyard, showing a different approach to local agentic AI. Really cool. Lightning is a 30B MoE with only 3B active parameters, distilled from Nemotron 3 Ultra and designed for the [image]
  • @lmstudio @lmstudio on x
    Nemotron 3.5 Lightning is available in LM Studio! The model is 30B MoE (3B active), can run very fast, and is trained for high volume agentic use cases. Model page: https://lmstudio.ai/...
  • @nvidiaai @nvidiaai on x
    Long-running agents spend most of their time executing: calling tools, validating results and delegating work. Nemotron 3.5 Lightning is built for this high-volume execution, at a size that can run anywhere from an NVIDIA DGX Spark to the data center. See it running agentic [vide…
  • @0xsero @0xsero on x
    I've been testing this model, excellent agent. With DSpark it's INCREDIBLY fast on a DGX Sparks 240+ tok/s Highly recommend trying this out. Not the best for coding but it's great at tool calling, exploration, apps, and being a general assistant. Best Nvidia model to date
  • @openrouter @openrouter on x
    NVIDIA Nemotron 3.5 Lightning is now live on OpenRouter. A 30B hybrid MoE with 3B active params, distilled from Nemotron 3 Ultra. Built for high-volume, specialized AI agent workloads, delivering up to 4× higher throughput and up to 30% faster task completion compared to similar
  • @nvidiaai @nvidiaai on x
    Introducing NVIDIA Nemotron 3.5 Lightning⚡ An open 30B MoE model with 3B active parameters, built for always-on agents to complete high-volume, specialized tasks faster. It delivers up to 4x the output speed of similar-sized models. [image]
  • @nvidiaai @nvidiaai on x
    Lightning is built to specialize. Post-train Nemotron 3.5 Lightning with NVIDIA NeMo for your domain data, tools, workflows and policies. Across cybersecurity, coding, legal and energy tasks, post-training improves accuracy for specialized work. [image]
  • @nvidiaai @nvidiaai on x
    Lightning pairs strong accuracy with speed. On PinchBench, it reaches 86% accuracy while completing 10,000 tasks 35% faster than Qwen3.6 35B at similar accuracy. [image]
  • @llm_wizard Chris on x
    🌩️Nemotron 3.5 Lightning is the newest member of the Nemotron 3 family - and the whole thing is that it's fassssst (and pretty smart too) - it's a 30B A3B model - runs on DGX Spark, ollama/llama.cpp support and all the good stuff. AND YOU KNOW IT'S AN OPEN LICENSE. Pushing [image…
  • @scaling01 @scaling01 on x
    “The largest Nemotron 4 model is expected to have at least 1 trillion parameters” lets fugging go, we need more huge open-weight models
  • @andrewcurran_ Andrew Curran on x
    NVIDIA is building its next-gen Nemotron 4 family to compete directly with leading Chinese open models and secure the open-weight crown for the U.S. The largest version will have at least 1 trillion parameters, according to original reporting from The Information. [image]
  • @paulharper.eurosky.social Paul Harper on bluesky
    Going from 550B parameters to 1T+ parameters is a significant doubling of my lack of interest...  [embedded post]
  • @aravsrinivas Aravind Srinivas on x
    A great American open weights MoE model that can run efficiently on your laptop or local hardware like the DGX Spark! You can use the larger Nemotron Ultra on Perplexity!
  • @liangsays Brent Liang on x
    thank you @nvidia for including @MTSlive under embargo with early access we've been playing with nemotron 3.5 lightning. very excited it's now out! [image]
  • @deryatr_ Derya Unutmaz on x
    Very excited that, after Meta released its open-source AI models yesterday, this morning another major American open-weights model was released by NVIDIA! Nemotron 3.5 lightning is a 30B-parameter model, 4x faster than similar sizes, with open weights, making it highly
  • @trajectorylabs @trajectorylabs on x
    Continual learning is a bet that the retraining loop will get cheaper over time. With larger models, you can maybe run this loop once every few weeks. But with smaller models, you can run it nightly, per customer. And it keeps recursing: a model per company, then a model per [ima…
  • @miaai_lab Mia on x
    Nvidia's Nemotron 3.5 Lightning is insanly fast on a DGX Spark, I'm blown away 🤯 [image]