A deep dive into Nvidia's Rubin CPX chip architecture, which is optimized for the prefill phase of inference, emphasizing compute FLOPS over memory bandwidth
the Rubin CTX is for inference, and it's a beast — sold as a rack as the unit, this one optimizes various LLM phases into silicon — notably: it does prefill in fp4 with low memory bandwidth and huge compute — semianalysis.com/2025/09/10/ a... Forums: r/AMD_Stock : Another Giant Leap: The Rubin CPX Specialized Accelerator & Rack r/NVDA_Stock : Another Giant Leap: The Rubin CPX Specialized Accelerator & Rack
SemiAnalysis
Context & Ripple Effects
Nvidia had already positioned Rubin CPX as a system for generative AI workloads, with availability planned for late 2026 in the earlier Rubin CPX launch announcement. This technical examination clarifies the design choice behind that positioning: prefill is being treated as a distinct silicon problem rather than a generic GPU task.
The architecture arrives amid a broader split between prefill and decode inference. Later coverage of a possible low-latency Rubin variant for decode-oriented reasoning underscores how workload phase is becoming a product-design boundary.
First-order effects
Nvidia’s customers get a rack-level inference option tuned for prefill, using FP4 and high compute throughput rather than maximizing memory bandwidth.
Rubin CPX makes phase-specific hardware a concrete deployment choice: operators can assess prefill capacity separately from the hardware serving other inference stages.
Second-order effects
Competing inference-chip vendors will face pressure to specify which inference phase their products optimize, rather than competing only on broad AI performance claims; Intel’s power- and cost-optimized inference GPU illustrates the adjacent positioning battle.
Rack-level specialization shifts evaluation toward end-to-end inference-system fit—software scheduling, interconnects, and workload mix—not just accelerator specifications.
Third-order effects
If prefill and decode continue to separate operationally, inference fleets are likely to become more heterogeneous, with different accelerators or system configurations allocated by model stage.
Value capture may move toward vendors that can package specialized silicon with the software and rack-level orchestration needed to keep those stages efficiently utilized.
The trend: AI inference is evolving from one-size-fits-all GPU deployment toward phase-specialized, system-level compute architectures.
@highyieldYT @nvidia Running some napkin math here, a perfect die with 16 GPCs has 32K CUDA + 256 RT + 1024 Tensor. The 6090 likely will follow a 10-15% binning so it could have: 27-29K CUDA + 218-230 RT + 872-920 Tensor So roughly a 28-35% performance uplift, assuming no uArch &…
@SemiAnalysis_ That's all? It's obvious you're biased against AMD. Just read the article. Is this suddenly a Grand Canyon Moat? Placing two different racks side by side and connecting them is the easiest thing to do. If this lateral step is Nvidia's move, it means they're falling…
NVIDIA has widened the gap for inference rack scale architecture yet again! Prefill specialized inference chips massively lower TCO per million input tokens on long context transformers As usual, other AI chip upstarts will follow this with prefill specialized chips, but later
The best thing about this is 3 year tco becomes substantially cheaper which would translate to better inference and per token economics. Oh and long context becomes a solved problem with this
Jensen the wizard “ AMD and custom silicon competitors may have made a small step forward in emulating Nvidia's 72-GPU rack scale design, but Nvidia has just made another Giant Leap, again leaving competitors very distant objects in the rear-view mirror.”
@KarbinC @nvidia So far we have only this image. It looks very close to GB202, but from the rumors we are not expecting a Rubin based gaming GPU. As you said, it might be really bad depending on the layout and units it has. But it looks like there's a raster engine and display ou…
Some quick assumptions about @Nvidia's new Rubin CPX. Quiet similar to GB202, with some obvious changes to the GPCs. Looks like it has full raster units (full GPU) with up to 256 ROPs. Could this be the RTX 6090? (even though Rubin was supposed to be AI only) [image]
HBM as a % total cost is increasing. Solution? Make a cheap, highly performant chip with no HBM for compute bound workflows Hard not to admire Nvidia's ingenuity with this one [image]
Rubin CPX is surprisingly close to a consumer chip, featuring units not required for HPC/AI: many GPCs, many ROPs, a lot of L2 cache, and even 4 display engines. However, it only has 192 SMs (same as GB202), so the RTX6090 get probably it's own chip. https://www.3dcenter.org/... …
NVIDIA's moat just got bigger — the Rubin CTX is for inference, and it's a beast — sold as a rack as the unit, this one optimizes various LLM phases into silicon — notably: it does prefill in fp4 with low memory bandwidth and huge compute — semianalysis.com/2025/09/10/ a.…