Nvidia says its Groq 3 LPX racks delivered 3,400 tokens per second in an Artificial Analysis benchmark running Gemma 4 31B with a 100,000-token input sequence
Nvidia's $20 billion bet on Groq's LPU tech sure looks like it was a good one. On Monday, the GPU giant offered the first glimpse …
Context & Ripple Effects
When Nvidia licensed Groq's inference technology for a reported $20B — against annual revenue sources pegged near $100M — the open question was whether the LPUs would ever ship as product. The March unveiling of the Groq 3 LPX rack promised 256 LPUs and 128GB of on-chip SRAM with H2 2026 availability; today Nvidia says the rack is in full production with Nebius as first customer, and is backing it with a published Artificial Analysis result.
That sequencing matters because Groq's independent route had stalled: the company cut its 2025 revenue projection to $500M+ citing data center capacity delays. A third-party throughput number on a real workload is the first external evidence that Nvidia's licensing bet converted into sellable hardware.
First-order effects
- Nebius, the first production customer, now has an independent Artificial Analysis datapoint to sell against — 3,400 tokens per second serving Gemma 4 31B with 100,000-token inputs, exactly the long-context regime where the rack's SRAM-first design differentiates.
- For Groq, shipping through Nvidia turns the LPU from a standalone cloud pitch into a volume product line inside the incumbent's stack, even as GroqCloud continues operating.
Second-order effects
- Because the Nvidia–Groq agreement is a non-exclusive license, other server vendors can build competing LPU racks, and the reported possibility of a Rubin SRAM variant for ultra-low-latency agentic workloads suggests Nvidia itself may fold the architecture into its own roadmap.
- Rivals selling GPU-based inference now face a named, published throughput reference point for 100K-token inputs, pressuring them to disclose comparable third-party-benchmarked numbers rather than internal figures.
Third-order effects
- If the pattern holds, inference silicon consolidates around SRAM-dense architectures absorbed by platform incumbents through licensing, leaving startups to monetize technology transfer rather than independent data center scale.
- Procurement shifts toward verified workload-level throughput from third-party benchmark houses like Artificial Analysis, making published long-context results the gate racks must pass before hyperscalers and neoclouds commit.
The trend: Inference hardware is becoming a benchmark-verified, workload-level market in which incumbents convert startup technology into shipping racks faster than those startups could scale alone.