/
Navigation
Chronicles
Browse all articles
Explore
Semantic exploration
Research
Entity momentum
Nexus
Correlations & relationships
Story Arc
Topic evolution
Drift Map
Semantic trajectory animation
Posts
Analysis & commentary
Pulse API
Tech news intelligence API
Browse
Entities
Companies, people, products, technologies
Domains
Browse by publication source
Handles
Browse by social media handle
Detection
Concept Search
Semantic similarity search
High Impact Stories
Top coverage by position
Sentiment Analysis
Positive/negative coverage
Anomaly Detection
Unusual coverage patterns
Analysis
Rivalry Report
Compare two entities head-to-head
Semantic Pivots
Narrative discontinuities
Crisis Response
Event recovery patterns
Connected
Search: /
Command: ⌘K
Embeddings: large
TEXXR

Chronicles

The story behind the story

days · browse · Enter similar · o open

A deep dive into Nvidia's Rubin CPX chip architecture, which is optimized for the prefill phase of inference, emphasizing compute FLOPS over memory bandwidth

the Rubin CTX is for inference, and it's a beast  —  sold as a rack as the unit, this one optimizes various LLM phases into silicon  —  notably: it does prefill in fp4 with low memory bandwidth and huge compute  —  semianalysis.com/2025/09/10/ a... Forums: r/AMD_Stock : Another Giant Leap: The Rubin CPX Specialized Accelerator & Rack r/NVDA_Stock : Another Giant Leap: The Rubin CPX Specialized Accelerator & Rack

SemiAnalysis

Context & Ripple Effects

Nvidia had already positioned Rubin CPX as a system for generative AI workloads, with availability planned for late 2026 in the earlier Rubin CPX launch announcement. This technical examination clarifies the design choice behind that positioning: prefill is being treated as a distinct silicon problem rather than a generic GPU task.

The architecture arrives amid a broader split between prefill and decode inference. Later coverage of a possible low-latency Rubin variant for decode-oriented reasoning underscores how workload phase is becoming a product-design boundary.

First-order effects

  • Nvidia’s customers get a rack-level inference option tuned for prefill, using FP4 and high compute throughput rather than maximizing memory bandwidth.
  • Rubin CPX makes phase-specific hardware a concrete deployment choice: operators can assess prefill capacity separately from the hardware serving other inference stages.

Second-order effects

  • Competing inference-chip vendors will face pressure to specify which inference phase their products optimize, rather than competing only on broad AI performance claims; Intel’s power- and cost-optimized inference GPU illustrates the adjacent positioning battle.
  • Rack-level specialization shifts evaluation toward end-to-end inference-system fit—software scheduling, interconnects, and workload mix—not just accelerator specifications.

Third-order effects

  • If prefill and decode continue to separate operationally, inference fleets are likely to become more heterogeneous, with different accelerators or system configurations allocated by model stage.
  • Value capture may move toward vendors that can package specialized silicon with the software and rack-level orchestration needed to keep those stages efficiently utilized.

The trend: AI inference is evolving from one-size-fits-all GPU deployment toward phase-specialized, system-level compute architectures.

Discussion

  • @rouanijihad @rouanijihad on x
    @highyieldYT @nvidia Running some napkin math here, a perfect die with 16 GPCs has 32K CUDA + 256 RT + 1024 Tensor. The 6090 likely will follow a 10-15% binning so it could have: 27-29K CUDA + 218-230 RT + 872-920 Tensor So roughly a 28-35% performance uplift, assuming no uArch &…
  • @speranzaguid GuidoSperanza on x
    @SemiAnalysis_ That's all? It's obvious you're biased against AMD. Just read the article. Is this suddenly a Grand Canyon Moat? Placing two different racks side by side and connecting them is the easiest thing to do. If this lateral step is Nvidia's move, it means they're falling…
  • @dylan522p Dylan Patel on x
    NVIDIA has widened the gap for inference rack scale architecture yet again! Prefill specialized inference chips massively lower TCO per million input tokens on long context transformers As usual, other AI chip upstarts will follow this with prefill specialized chips, but later
  • @basedjensen @basedjensen on x
    The best thing about this is 3 year tco becomes substantially cheaper which would translate to better inference and per token economics. Oh and long context becomes a solved problem with this
  • @apples_jimmy @apples_jimmy on x
    Jensen the wizard “ AMD and custom silicon competitors may have made a small step forward in emulating Nvidia's 72-GPU rack scale design, but Nvidia has just made another Giant Leap, again leaving competitors very distant objects in the rear-view mirror.”
  • @highyieldyt @highyieldyt on x
    @KarbinC @nvidia So far we have only this image. It looks very close to GB202, but from the rumors we are not expecting a Rubin based gaming GPU. As you said, it might be really bad depending on the layout and units it has. But it looks like there's a raster engine and display ou…
  • @highyieldyt @highyieldyt on x
    Some quick assumptions about @Nvidia's new Rubin CPX. Quiet similar to GB202, with some obvious changes to the GPCs. Looks like it has full raster units (full GPU) with up to 256 ROPs. Could this be the RTX 6090? (even though Rubin was supposed to be AI only) [image]
  • @aj_kourabi @aj_kourabi on x
    HBM as a % total cost is increasing. Solution? Make a cheap, highly performant chip with no HBM for compute bound workflows Hard not to admire Nvidia's ingenuity with this one [image]
  • @3dcenter_org @3dcenter_org on x
    Rubin CPX is surprisingly close to a consumer chip, featuring units not required for HPC/AI: many GPCs, many ROPs, a lot of L2 cache, and even 4 display engines. However, it only has 192 SMs (same as GB202), so the RTX6090 get probably it's own chip. https://www.3dcenter.org/... …
  • @timkellogg.me Tim Kellogg on bluesky
    NVIDIA's moat just got bigger  —  the Rubin CTX is for inference, and it's a beast  —  sold as a rack as the unit, this one optimizes various LLM phases into silicon  —  notably: it does prefill in fp4 with low memory bandwidth and huge compute  —  semianalysis.com/2025/09/10/ a.…
  • r/AMD_Stock r on reddit
    Another Giant Leap: The Rubin CPX Specialized Accelerator & Rack
  • r/NVDA_Stock r on reddit
    Another Giant Leap: The Rubin CPX Specialized Accelerator & Rack