In August 2026, Nvidia named Nebius the first customer for a rack containing 256 Groq inference chips. Nvidia had licensed Groq’s design, yet Groq remained independent and planned to build its own capacity.

Key takeaways

  • Nvidia’s Groq 3 LPX rack contains 256 Groq 3 LPUs, 128GB of on-chip SRAM and 40 PBps of SRAM bandwidth.
  • Nvidia named Nebius as LPX’s first customer in August 2026.
  • Groq raised $650 million after the Nvidia deal and set a target of 200MW of capacity by the end of 2027.
  • The reported Nvidia–Groq agreement was a roughly $20 billion nonexclusive license, not an announced acquisition.
  • Nvidia reported 3,400 tokens per second for LPX on Gemma 4 31B with a 100,000-token input sequence.

Different serving jobs reward different silicon

In March 2025, the Financial Times reported Barclays’ projection that inference capital spending would reach $208.2 billion in 2026 and surpass training spending within two years. That was a forecast, not a measured 2026 total. It helps explain why chipmakers began designing around particular serving jobs rather than treating every inference request as a smaller training workload.

Prefill processes an incoming prompt; decode produces the response. SemiAnalysis described Nvidia’s Rubin CPX as optimized for prefill, favoring compute throughput over memory bandwidth. Groq’s SRAM-heavy approach makes a different trade-off. Model design also changes the hardware requirement: Google’s Gemma 3 and Cohere’s Command A were reported to run on one or two H100s. A buyer choosing serving hardware must specify the model, prompt and latency target before a chip-speed claim becomes useful.

Nvidia licensed the design and brought in its builders

The Information reported a roughly $20 billion Nvidia–Groq licensing deal in December 2025. Groq described the agreement as nonexclusive and said it would keep operating independently. CEO Jonathan Ross and other senior executives joined Nvidia. Nvidia thus gained access to both an inference architecture and people who knew how to turn it into a product, without announcing an acquisition of Groq.

In March 2026, Nvidia announced the Groq 3 LPX: a rack containing 256 Groq 3 LPUs, 128GB of on-chip SRAM and 40 PBps of SRAM bandwidth. Nvidia positioned it alongside its Vera Rubin NVL72 rack, not as a replacement for the GPU platform. The unit Nvidia presented to customers was a configured system bearing its own name.

Groq’s delays put delivery ahead of speed

Groq had already encountered the difference between a fast design and deliverable capacity. In July 2025, The Information reported that Groq had cut its revenue projection for that year from more than $2 billion to more than $500 million, citing data-center-capacity delays. Neither projection was actual revenue, but the revision identified a constraint chip performance alone could not lift.

In August 2026, Nvidia said LPX had entered full production and named Nebius its first customer. Nvidia also reported 3,400 tokens per second in an Artificial Analysis benchmark using Gemma 4 31B and a 100,000-token input sequence. Nvidia reported that result for one model and prompt without establishing LPX’s cost per token, power use or utilization against a GPU deployment. Nor does the published record establish who sets LPX prices, allocates racks or qualifies customers.

Other builders are choosing different routes to serving

AMD acquired Taalas, whose design integrates model weights directly into silicon; early demonstrations were reported at up to 17,000 tokens per second. OpenAI instead previewed a Cerebras-powered Ultrafast API tier, reporting up to 750 output tokens per second for GPT-5.6 Sol. AMD acquired a design, OpenAI packaged specialist hardware as a service, and Nvidia put licensed LPUs in a rack. Those speed figures describe different workloads and cannot rank the three offerings on delivered economics.

Specialization also has a limit. An analysis of fully agentic workloads argued that applications without a person waiting for each response may place less value on ultra-low latency. A separate tokens-per-dollar comparison favored Nvidia over the TPU v6e and AMD MI300X on its chosen inference metric. Neither finding rules out specialist chips; both make the buyer’s workload and operating cost more consequential than the fastest published output rate.

“Nonexclusive” leaves the sales boundary open

Groq raised $650 million after the Nvidia deal to expand its own data-center business and said it aimed for 200MW of capacity by the end of 2027. That planned capacity gives Groq a potential route to customers separate from LPX. It is evidence against treating the license as complete commercial absorption, though a financing round and a target do not establish how much capacity Groq can deliver.

The Justice Department was reportedly examining whether Nvidia sought to sidestep antitrust scrutiny in the Groq transaction. An inquiry is not a finding. It does make the undisclosed commercial terms consequential: customers and competitors cannot infer from the word “nonexclusive” who controls LPX supply, pricing or access to Groq-based inference through each company.

Frequently asked questions

When was the Nvidia Groq 3 LPX expected to become available?

CRN’s March 2026 coverage said the LPX was available in the second half of 2026. Nvidia subsequently said it had entered full production in August 2026.

What does Groq’s 200MW target measure?

It is a stated data-center power-capacity target, not a chip count or a guaranteed customer-throughput figure. Groq said it aimed to reach that capacity by the end of 2027.

Did Nvidia buy Groq outright?

No acquisition was announced in the piece’s account. Groq described the arrangement as nonexclusive, said it would remain independent, and said GroqCloud would continue operating.

What remains unknown about buying or operating an LPX system?

The published record does not disclose LPX pricing, rack allocation, customer qualification, cost per token, power use, or utilization relative to GPU deployments. Those omissions prevent a delivered-economics comparison from the reported speed benchmark alone.

Key milestones in the Groq–Nvidia commercialization path

  • March 2025 — The Financial Times reported Barclays’ forecast that inference capital spending would reach $208.2 billion in 2026.
  • July 2025 — The Information reported that Groq cut its 2025 revenue projection from more than $2 billion to more than $500 million, citing data-center-capacity delays.
  • December 2025 — The Information reported a roughly $20 billion Nvidia–Groq licensing deal.
  • March 2026 — Nvidia announced the Groq 3 LPX rack with 256 Groq 3 LPUs, 128GB of on-chip SRAM and 40 PBps of SRAM bandwidth.
  • August 2026 — Nvidia said LPX had entered full production and named Nebius as its first customer.
  • End of 2027 — Groq said it aimed for 200MW of its own data-center capacity.

If Groq can finance and sell its own serving capacity, the license leaves customers two routes to its architecture. If Nvidia alone can supply and support LPX at scale, customers will encounter that architecture chiefly through Nvidia. The purchase order, rather than the license label, marks the boundary.