On August 11, 2026, Nvidia released Nemotron 3.5 Lightning, an open AI model it says can generate output up to four times faster, and NeMo Switchyard, software that chooses which model handles each request. Reports at the same time tied Nvidia, Apollo, BlackRock and others to a $500 billion infrastructure package. The chip supplier was reaching for both request routing and the financing that determines which capacity gets built.

Key takeaways

  • Nvidia released Nemotron 3.5 Lightning and NeMo Switchyard on August 11, 2026.
  • Nemotron 3.5 Lightning is a 30B-parameter mixture-of-experts model.
  • IBM and Together AI signed a confirmed $240 million multi-year agreement for an IBM Cloud inference cluster using Nvidia HGX B300 systems.
  • The reported $500 billion AI-infrastructure package involving Nvidia, Apollo, BlackRock and others was still under discussion, with no commitments or deployment announced.
  • AI infrastructure companies borrowed more than $100 billion in 2025.

CUDA gives Nvidia the technical attachment point for that expansion. Open models, routing software, reference clusters and capital alliances answer the same shift: heterogeneous, cost-sensitive inference makes choosing among resources more valuable than adding undifferentiated execution.

The release paired a 30B-parameter mixture-of-experts model with an agentic router. Nemotron gives Switchyard one destination Nvidia controls; the router can send work elsewhere when latency, quality or cost demands it.

In the 2024 comparison period, 15.2% of Nvidia articles used enterprise framing; in 2026, 23.5% did. Consumer framing fell from 18.3% to 10.9%. Coverage followed Nvidia up the stack, from devices and chips toward deployment and capital allocation.

CUDA turned compatibility into accumulated leverage

Nvidia spent years pulling workloads toward its hardware through software. In 2018, the company added GPU support to Kubernetes and contributed its enhancements to the open-source orchestration community. Nvidia also launched Rapids with IBM, HPE and other partners, tying GPU acceleration to data-science and machine-learning workflows. In 2019, Nvidia said it would bring CUDA acceleration to Arm CPUs and share its AI and high-performance-computing software with Arm.

Those integrations widened the range of workloads builders could bring onto Nvidia systems. Organizations then accumulated code, tooling and operating practices around CUDA, making it cheaper to stay than to rewrite.

Chinese AI labs show what that accumulation looks like after years of use. Sources say those labs still train leading language models on Nvidia chips because moving from CUDA to Huawei’s CANN requires major code changes. The hardware purchase created the initial relationship; software preserved it across subsequent purchasing decisions.

Those labs remain attached to CUDA, but migration costs do not give Nvidia authority over their models, routing policies or service-quality thresholds. Nvidia enters the next contest with a powerful advantage, not a guaranteed victory.

The expensive decision now occurs before execution

Documents shown to investors reportedly indicate that OpenAI and Anthropic spend more than half of their revenue on inference. At that scale, every request forces a decision about how much quality, latency and computation the service can afford.

Reported inference costs as a share of revenue at OpenAI and Anthropic

Nemotron gives Switchyard one candidate inside that decision. An operator can direct different tasks toward different models rather than assign every request the same system and cost profile. Nvidia’s claimed fourfold output gain would strengthen that proposition, though independent testing has not established the figure.

Operators must balance portability, latency, cost, resilience and governance, while the router adds another service whose policies can fail. A routing error can erase savings; an opaque policy can complicate procurement and accountability. Raw throughput captures only one part of the buyer’s cost.

Other builders have moved toward the same pressure point. Inferact, founded by the creators of the vLLM inference engine, raised a $150 million seed round at an $800 million valuation to commercialize inference serving. OpenAI engineers, meanwhile, reportedly found an internal method that could more than halve inference cost, allowing OpenAI to capture that gain without adopting Nvidia’s router.

Nvidia, OpenAI and Inferact converged on serving costs from different directions. Their approaches establish the value of optimization while leaving ownership of that value unsettled.

Open models make reference systems more valuable

Operators can standardize infrastructure even as they change models. IBM and Together AI signed a confirmed $240 million, multi-year agreement to build an inference cluster on IBM Cloud using Nvidia HGX B300 systems for open-source models. Together AI can support a changing model catalog while IBM standardizes the environment beneath it.

Nvidia has cultivated this implementation path for years. Its 2018 HGX-2 combined 16 GPUs into a platform for AI and high-performance computing. The Tesla T4 targeted data-center inference, with Nvidia claiming up to twelve times the performance of the prior Tesla P4 at the same power. HGX B300 systems extend that platform logic into a market where operators increasingly care about the complete serving system.

Cloud operators can replace or add open models without rebuilding the cluster architecture each time. Nvidia can prosper even if no single model provider dominates, provided many providers converge on its reference systems, software libraries and operating practices.

The IBM–Together AI agreement gives operators a concrete procurement path. They can buy a multi-year service for open inference without assembling every hardware and software component independently. Open models let buyers change catalog entries while a standardized implementation preserves the Nvidia environment underneath.

Financing decides which capacity exists

Nvidia, Apollo, BlackRock, Blackstone, Brookfield, Goldman Sachs, KKR and others were reportedly working on a $500 billion funding package, moving allocation further upstream through AI infrastructure finance. The package was still under discussion, before commitments or deployment. Its participants were targeting the financing constraint on new capacity.

Other transactions expose the mechanism at smaller scale. Lambda was reportedly marketing a $917 million leveraged loan to finance GPUs under a contract with Nvidia. Nvidia reportedly agreed to invest $2 billion in power-infrastructure developer Lancium, with another $1 billion tied to performance thresholds at a Stargate campus. The chip order, power site, customer contract and loan increasingly arrive as one financial object.

AI infrastructure companies borrowed more than $100 billion in 2025, while debt investors charged smaller companies higher interest rates as they questioned unproven AI businesses. Lenders can accelerate construction, but their carrying costs must be supported by inference revenue. A cluster financed with interest still has to sustain a business model.

Lenders and asset owners allocate capacity by deciding which contracts are bankable, which power projects receive capital and which hardware configurations sit behind the collateral. If they prefer Nvidia-backed deployments because customers, software and resale markets already recognize them, financing can reinforce the installed base without formal coordination among buyers.

Nvidia benefits when institutional investors fund systems that consume its products, while lenders and asset owners absorb more exposure if AI revenue fails to cover the buildout. Those investors create demand only while they accept the utilization assumptions beneath it; underwriting, rather than the headline sum, decides how much capacity gets built.

Routing power has to be earned on every request

Rival chipmakers can attack the same allocation problem from a different boundary. AMD, Nvidia’s principal AI-processor challenger, acquired Taalas, a startup that integrates model weights directly into silicon with the promise of improving inference performance. Taalas compresses the distance between model and hardware instead of inserting a broader routing layer above them.

Customers can also retain optimization inside their own systems, as OpenAI’s reported cost reduction demonstrates. Cloud providers, model hosts and enterprises will benchmark a router on total operating cost, latency, reliability, portability and policy control. They can keep workloads on internal systems wherever Switchyard loses that comparison.

Borrowed money tightens the test because utilization shortfalls leave financed capacity carrying interest. A routing layer that improves hardware use can support the economics. If it mainly preserves hardware attachment, it becomes another cost center in an already leveraged system.

Frequently asked questions

What evidence would substantiate Nvidia’s claim of up to 4× faster output from Nemotron 3.5 Lightning?

The piece says independent testing has not established the claimed gain. It does not provide benchmark methodology, workloads, baseline models or a schedule for third-party evaluation.

Which non-Nvidia models or hardware platforms can NeMo Switchyard route work to?

The article says Switchyard can send work elsewhere when latency, quality or cost requires it, but it does not identify supported external models, clouds or hardware platforms.

Which firms have actually committed capital to the reported $500 billion infrastructure package?

No commitments are identified. The reporting describes Nvidia, Apollo, BlackRock, Blackstone, Brookfield, Goldman Sachs, KKR and others as working on a package that remained under discussion.

What performance thresholds would unlock the additional $1 billion in Nvidia’s reported Lancium investment?

The piece states that $1 billion was tied to performance thresholds at a Stargate campus, but it does not disclose the thresholds, measurement criteria or payment timing.

Reported AI-infrastructure financing amounts

ArrangementAmountStatus or condition
Infrastructure package involving Nvidia, Apollo, BlackRock and others$500 billionReportedly under discussion; no commitments or deployment announced
Lambda leveraged loan for GPUs under a contract with Nvidia$917 millionReportedly being marketed
Nvidia investment in power-infrastructure developer Lancium$2 billionReportedly agreed
Additional Nvidia investment in Lancium$1 billionTied to performance thresholds at a Stargate campus
AI infrastructure-company borrowing in 2025More than $100 billionReported aggregate borrowing

Nvidia keeps the dispatch desk only while customers can verify that Switchyard lowers the cost of useful work and leaves them control over models and policy. CUDA made Nvidia expensive to leave at compile time; every request gives the customer another chance to leave at run time. The $500 billion may finance the destinations, but it cannot decide where the next request goes.