In March 2026, Nvidia announced the Nvidia Groq 3 LPX, a rack containing 256 Groq LPUs. By August, the system had entered full production. Ten years earlier, Nvidia had framed the Tesla P100 around more than 15 billion transistors and 16GB of high-bandwidth memory; now it was shipping a would-be GPU substitute inside an Nvidia rack.
Key takeaways
- Nvidia and Groq confirmed a $20 billion non-exclusive licensing agreement related to AI inference.
- The Nvidia Groq 3 LPX rack contains 256 Groq LPUs and pairs them with 128GB of on-chip SRAM.
- Nvidia reported 3,400 tokens per second for Groq 3 LPX on an Artificial Analysis benchmark using Gemma 4 31B and a 100,000-token input sequence.
- Investor materials reviewed by The Wall Street Journal showed both OpenAI and Anthropic reporting inference costs above half of revenue.
- Nvidia disclosed $3.5 billion in guarantees for businesses leasing land, power and data-center facilities—four times its third-quarter level.
Nvidia is building its next moat around the commercial terms and infrastructure surrounding heterogeneous AI compute: integration, delivery, financing, and utilization. It is licensing specialist architectures, turning them into Nvidia-branded systems, and underwriting parts of deployment. As inference becomes the recurring cost center, control shifts toward the supplier that can package the right latency, throughput, memory, and cost profile into capacity a customer can buy and keep utilized.
Nvidia bought option value instead of demanding exclusivity
Nvidia and Groq entered a confirmed $20 billion non-exclusive licensing agreement related to AI inference. Groq founder Jonathan Ross, who helped create Google’s TPU, joined Nvidia, along with several senior Groq executives. Nvidia secured access to the architecture and the people who knew how to turn it into a system.
Nvidia turned the agreement into a product. The Groq 3 LPX pairs its LPUs with 128GB of on-chip SRAM and sits alongside the GPU-based Vera Rubin NVL72. Nebius became its first customer when full production began in August.
Groq retained real independence. The non-exclusive agreement gives Nvidia no established right to prevent Groq-derived technology from serving other buyers. The Justice Department is also reportedly examining whether the arrangement skirted antitrust scrutiny.
Nvidia avoided a binary wager. A full acquisition would have concentrated regulatory and technical risk; ignoring Groq would have left Nvidia exposed if LPUs won economically important serving jobs. By licensing the technology, hiring its leaders, and integrating it into a product, Nvidia gained a shipping option without taking exclusive control. Groq can remain outside Nvidia while its architecture sits inside an Nvidia rack, making the purchase order more consequential than the corporate boundary.
Different inference jobs reward different silicon
Nvidia’s own roadmap shows why inference resists a single architectural answer. The company designed Rubin CPX for the prefill phase of inference and emphasized compute FLOPS over memory bandwidth. Groq 3 LPX follows another route, pairing 128GB of SRAM with 40 petabytes per second of SRAM bandwidth.
Nvidia reported 3,400 tokens per second for Groq 3 LPX in an Artificial Analysis benchmark running Gemma 4 31B with a 100,000-token input sequence. The result shows how a narrow architecture performs under a defined model, context, and latency profile.
Nuance Labs shows why applications split the market. The company raised a $50 million Series A, with Nvidia participating, to build low-latency conversational avatars. In a face-to-face avatar, response delay shapes the product experience. A background agent can exchange speed for lower cost or higher aggregate throughput because no human waits on each response.
GPUs remain formidable in inference. In one comparison of Nvidia, Google, and AMD hardware, Nvidia achieved roughly five times the tokens per dollar of Google’s TPU v6e and twice that of AMD’s MI300X. Nvidia also held an estimated 95% share of machine-learning GPUs in 2023, giving it a large installed base and deep developer familiarity.
A specialist must beat more than a GPU benchmark. It must overcome software maturity, deployment habits, and the cost of introducing another architecture into production. Nvidia treated Groq as commercially useful for selected jobs while its flagship architecture still led broad deployments.
Model margins now choose the serving stack
OpenAI and Anthropic have put the serving decision directly onto the income statement. Investor materials reviewed by The Wall Street Journal showed both companies reporting inference costs exceeding half of revenue. OpenAI engineers later reportedly told colleagues that they had found a way to more than halve inference cost.
For a model provider spending more than half its revenue on inference, improvements in model efficiency, batching, memory use, utilization, or hardware selection flow directly to gross margin. The provider must buy for cost per useful output under actual service requirements, with peak performance as only one input.
At that cost level, the CFO and the infrastructure engineer are choosing the same serving stack.
Barclays projected 2026 inference capital expenditure at $208.2 billion and expected inference spending to surpass training within two years. General Compute then obtained a $400 million loan backed by inference-specific chips, described as the first financing of its kind.
By accepting inference-specific chips as collateral, Upper90 had to assess utilization, resale demand, technical obsolescence, and the contracts attached to the machines. Financing AI inference infrastructure turns architectural fitness into an underwriting question. A chip that performs well but cannot stay busy makes weak collateral.
Specialists are attracting money before standards settle
The creators of the open-source vLLM inference engine founded Inferact to commercialize that layer and raised a $150 million seed round. By building a company around serving software, they separated inference optimization from model development.
Netherlands-based Euclyd raised more than €200 million in a Series A for inference chips. Cerebras reported second-quarter revenue of $180 million, up 74% year over year, and raised its annual revenue and gross-margin forecasts. DeepSeek is reportedly developing an inference chip to reduce its reliance on Nvidia and Huawei hardware.
The vLLM founders, Euclyd’s investors, Cerebras customers, and DeepSeek engineers each faced their own version of recurring serving cost. Each isolated a layer where specialized hardware or software might capture the savings.
Google’s Gemma 3 and Cohere’s Command A were reported to run on one or two H100s, reducing the hardware required for an individual deployment. More efficient models can shrink demand per task even as usage grows. As each deployment needs fewer accelerators, suppliers have less room to hide an architecture that fits the workload poorly.
Nvidia is pricing utilization risk into the rack
Nvidia is also using its balance sheet to reduce deployment risk. The company disclosed $3.5 billion in guarantees to businesses leasing land, power, and data-center facilities, four times its third-quarter level. Those guarantees can help infrastructure companies without strong credit ratings secure the inputs required to install capacity.
CoreWeave pioneered the use of GPU-backed debt and special-purpose vehicles. Nvidia reportedly offered to rent unused GPUs from young cloud providers in exchange for a share of their revenue. The company also sought to connect GPU holders with Nordic data-center operators that had available deployment capacity.
Cloud operators carry construction and operating exposure. Lenders price the collateral, and customers commit to capacity. Nvidia uses guarantees, rental backstops, and matchmaking to reduce the chance that purchased hardware sits idle. By posting $3.5 billion in guarantees, Nvidia has put part of its own credit behind the buildout.
Frequently asked questions
Can the reported 3,400-token-per-second result be directly compared with GPU performance claims?
Not from the information provided. The Groq figure is tied to Gemma 4 31B and a 100,000-token input sequence; a like-for-like comparison would also need the same model, context length, latency target, batching configuration and cost basis on competing hardware.
How much Groq 3 LPX capacity has Nebius committed to buy?
The piece identifies Nebius as the first customer when full production began, but it does not disclose an order size, contract value, deployment schedule or utilization commitment.
Which companies receive Nvidia's $3.5 billion in guarantees, and on what terms?
The piece says the guarantees support businesses leasing land, power and data-center facilities, but it does not name all beneficiaries or disclose durations, fees, collateral requirements or default conditions.
How much of Barclays' projected $208.2 billion in 2026 inference capex will flow to Nvidia or Groq-based systems?
The projection is for aggregate inference capital expenditure. The piece provides no vendor-by-vendor allocation, so it cannot establish Nvidia's or Groq LPX's share of that spending.
Groq 3 LPX path to production
- March 2026 — Nvidia announced the Groq 3 LPX, a rack containing 256 Groq LPUs.
- August 2026 — The Groq 3 LPX entered full production, with Nebius as its first customer.
A cloud buyer now compares GPU and LPU capacity alongside delivery schedules, financing terms, integration work, and utilization support. Nvidia can remain the counterparty when a specialist chip best fits the job. The 256 Groq LPUs inside its rack show the bargain: Groq can win the workload while Nvidia keeps the purchase order.