Groq cut its 2025 revenue projection from more than $2 billion to more than $500 million, citing delays in data-center capacity. The company had a fast inference accelerator and customers to serve, but not enough deployed infrastructure to convert them into revenue. Nvidia later entered a reported $20 billion licensing agreement and put 256 Groq 3 LPUs into one production rack. Groq’s chip advantage now depended on financing and deployment.
Key takeaways
- Barclays projected $208.2 billion of inference capital expenditure for 2026.
- Groq set a 200-megawatt capacity target for the end of 2027.
- General Compute obtained a $400 million loan backed by inference-specific chips.
- Google priced Gemini 3.7 Flash at $0.75 per million input tokens and $3.75 per million output tokens through 2026.
- OpenAI’s Cerebras-powered Ultrafast API tier claimed up to 750 output tokens per second for GPT-5.6 Sol.
Reporters first described Groq in 2017 as a stealth startup founded by Google Tensor Processing Unit engineers that had raised $10.3 million. Between that round and Nvidia’s rack, inference split into technical jobs, capacity commitments and routes to customers.
Different inference jobs reward different silicon
A model server handles several distinct tasks. Prefill, decode, long-context processing and latency-sensitive interaction impose different constraints, so builders have started assigning architectures to phases rather than asking one accelerator to win every benchmark.
Nvidia’s Rubin CPX illustrates that decomposition. Nvidia designed the chip specifically for prefill and prioritized compute FLOPS over memory bandwidth. Groq attacks another part of the workload with large amounts of on-chip SRAM and an execution model aimed at fast, predictable generation. OpenAI and Broadcom made the same structural move from the buyer side: they developed the LLM-optimized Jalapeño inference chip from design through manufacturing tape-out in nine months.
Working independently, model operators and chip designers encountered different bottlenecks inside the same inference request, then optimized the phase where general-purpose hardware imposed the largest penalty.
Nvidia subsequently packaged Groq’s architecture as the Groq 3 LPX rack, containing 256 Groq 3 LPUs, 128 GB of on-chip SRAM and 40 petabytes per second of SRAM bandwidth. Nvidia later reported 3,400 tokens per second in an Artificial Analysis benchmark running Gemma 4 31B with a 100,000-token input sequence.
That reported benchmark supports a workload advantage while leaving rack utilization, customer pricing, energy costs, software migration and model coverage unanswered. Benchmarks show how fast the engine moves; infrastructure economics tracks whether an operator can keep the seats full.
Incumbents can respond to workload fragmentation by placing a specialist beside CPUs, GPUs, networking and software, then routing each phase to the appropriate device. The specialist keeps its technical edge inside the incumbent’s commercial system.
Capacity turns chip companies into infrastructure operators
Groq’s forecast cut exposed the gap between accelerator supply and serving capacity. Without deployed racks, Groq could not turn customer demand into revenue.
Groq later raised $650 million to expand capacity. Its end-of-2027 power target sat inside a much larger spending shift: Barclays projected that inference capital expenditure would surpass training expenditure within two years. Investors were financing the processor together with the sites, power and systems required to serve tokens.
General Compute received a $400 million loan backed by inference-specific chips, an early case of specialist hardware serving as collateral. A deep resale market for those chips remains unproven, yet lenders have started underwriting inference capacity as infrastructure rather than venture inventory with a power cord.
Vendors create capacity lag when they commit money before sites come online, then wait for customer qualification and utilization to reveal whether they sized the deployment correctly. They pay for errors in either direction: idle racks damage returns, while missing capacity strands demand and postpones revenue.
Frontier labs complicate underwriting because large buyers increasingly obtain capacity from suppliers that also provide financing. In the broader frontier-lab financing loop, chipmakers and cloud operators become capital partners, committing to deployments before customers have generated the revenue needed to absorb the hardware.
Nvidia absorbed differentiation without buying the company
Nvidia used licensing, hiring and productization rather than a conventional acquisition. Under the non-exclusive agreement, Nvidia obtained rights to Groq’s inference technology, while CEO Jonathan Ross and other senior executives joined Nvidia. Sources later said roughly 90% of Groq employees would move to Nvidia and that most shareholders would receive payouts tied to the $20 billion valuation.
Groq said it would continue operating independently and subsequently raised capital for its planned cloud expansion. Nvidia gained the architecture and much of the organization capable of developing it. The non-exclusive terms preserved an independent corporate entity, while talent and product integration let Nvidia capture the technology inside its own systems portfolio.
Nvidia then converted the license into capacity that customers could buy. The company put Groq 3 LPX into full production, and Nebius became the first announced customer. The rack placed the LPU inside Nvidia’s broader data-center product machine.
Nvidia and Groq did not disclose who would control LPX customer qualification, allocation, pricing or cloud access. Whoever controls that route to deployment can capture value even when the underlying architecture remains available elsewhere.
The Department of Justice is reportedly examining whether Nvidia used the licensing structure to avoid antitrust scrutiny. The inquiry does not establish wrongdoing, but it raises the transaction cost of using licenses plus mass hiring as a substitute for acquisition. Regulators can constrain how an incumbent internalizes a challenger even when they cannot erase the technical logic for integration.
Groq alone cannot establish that every specialist will be absorbed. Qualcomm’s reported talks to acquire Tenstorrent for $8 billion to $10 billion show strategic interest, not a completed pattern. Groq’s post-agreement capital raise also preserves a genuine independent-capacity experiment. Specialists and incumbents are testing both structures at once: specialist-operated infrastructure and specialist technology inside incumbent distribution.
Falling token prices shorten the amortization clock
Dedicated inference capacity must recover its cost while the price of serving continues to move. OpenAI and Anthropic projected that inference costs exceeded half of revenue, giving both companies a powerful incentive to cut the bill through models, software, silicon and procurement.
OpenAI engineers reportedly found a method that could more than halve inference costs. Google priced Gemini 3.7 Flash at $0.75 per million input tokens and $3.75 per million output tokens through 2026. Google’s price does not reveal its underlying cost, but it sets a market reference that competing providers must answer.
Model and software improvements can move faster than data-center construction. A rack financed against today’s token economics may enter service after a provider reduces the compute required per request or cuts its price to gain volume. Dedicated hardware can remain technically excellent while its owner earns a poor return if utilization fails to offset that decline.
Agentic workloads complicate the low-latency case further. Human-facing applications reward immediate output because a person waits on each response. For autonomous workloads with no person waiting at every intermediate step, throughput, reliability or cost per useful task can matter more than peak tokens per second.
Interactive applications can justify a premium for Groq-style latency, while batch-like agent execution rewards another architecture or a cheaper serving tier. Capacity underwriters therefore need a workload mix, not merely a benchmark, because each workload assigns a different price to speed.
Buyers keep specialists alive by refusing a single stack
Model builders still have reasons to maintain credible alternatives. OpenAI reportedly sought inference chips from AMD, Cerebras and Groq after becoming dissatisfied with some Nvidia products. The company then introduced an Ultrafast API tier powered by Cerebras, claiming up to 750 output tokens per second and as much as a 14-fold speed increase for GPT-5.6 Sol. By launching a live API tier, OpenAI created independent demand for specialist infrastructure.
Other buyers and suppliers are building additional options. AMD acquired Taalas, which integrates model weights directly into silicon and produced early demonstrations of up to 17,000 tokens per second. ByteDance and InnoStar reportedly began developing a low-cost inference chip modeled after Groq’s LPUs.
Buyers get three forms of value from these alternatives even when they do not displace the market leader: workload-specific performance, resilience against constrained supply and bargaining leverage over the incumbent. They can preserve that leverage only if they can move models and workloads without prohibitive software or operational costs. A technically superior chip with no practical migration path is a negotiating anecdote.
Incumbents can spread rack integration, software support and distribution costs across a larger installed base. Specialists retain room where their architecture creates a large enough workload advantage to overcome that integration discount. They gain more room when capital providers finance independent capacity and major buyers commit enough demand to fill it.
Frequently asked questions
When was Nvidia Groq 3 LPX announced, and when was it expected to become available?
Nvidia announced the 256-LPU Groq 3 LPX rack on March 16–17, 2026. The announcement said it would be available in the second half of 2026.
Can Nvidia’s reported 3,400 tokens per second be directly compared with OpenAI’s 750-token-per-second Cerebras claim?
Not reliably. Nvidia’s figure was reported for Gemma 4 31B with a 100,000-token input sequence, while OpenAI’s claim concerned GPT-5.6 Sol; model, prompt length and serving conditions differ.
How much funding had Groq raised before its later capacity-expansion round?
Sources cited in the evidence said Groq had raised approximately $1.8 billion to date. That figure is reported rather than independently detailed in the piece.
What is the comparable specialist-chip transaction to watch beyond Groq?
Qualcomm was reportedly in talks to acquire Tenstorrent for $8 billion to $10 billion. The discussions were reported talks, not a completed acquisition.
From Groq startup to rack-capacity targets
- 2017 — Groq was described as a stealth startup founded by Google TPU engineers that had raised $10.3 million.
- 2025 — Groq cut its 2025 revenue projection from more than $2 billion to more than $500 million, citing data-center-capacity delays.
- March 16–17, 2026 — Nvidia announced the Groq 3 LPX inference rack with 256 Groq 3 LPUs, 128 GB of on-chip SRAM and 40 PBps of SRAM bandwidth.
- H2 2026 — The announced availability window for Nvidia Groq 3 LPX.
- End of 2027 — Groq’s stated capacity target is 200 MW.
Groq’s 2025 forecast cut began with racks that arrived late. Nvidia’s 256-LPU rack closes that loop: the LPU survives, but the name on the rack identifies who controls deployment—and who must turn its engineering gain into cheaper tokens.