Nvidia announces plans to upgrade its AI accelerators annually, a Blackwell Ultra chip for 2025, and Blackwell's successor, Rubin, for 2026 that will use HBM4
The explicit HBM4 reference matters because it connects Nvidia’s accelerator roadmap to the memory supply chain and gives infrastructure buyers a clearer view of when another platform transition is expected.
First-order effects
Nvidia gives cloud providers, system builders, and enterprise buyers a stated annual planning cycle for accelerator deployments, with an intermediate Blackwell Ultra step before Rubin.
Rubin’s planned use of HBM4 makes next-generation high-bandwidth memory a design dependency for Nvidia’s 2026 platform rather than an unspecified future upgrade.
Second-order effects
HBM suppliers and server-system partners gain a clearer signal to align product qualification, capacity plans, and platform integration around Nvidia’s cadence.
Rival AI-chip vendors face pressure to show not only peak performance but credible, repeatable roadmaps that reduce buyers’ risk of committing to a competing platform.
Third-order effects
If Nvidia sustains the cadence, AI infrastructure purchasing may increasingly resemble a rolling platform-refresh cycle, with compute, memory, networking, and systems decisions made together.
The roadmap reinforces the importance of an annual accelerator release schedule: platform control can compound when customers synchronize procurement and software optimization to one vendor’s successive generations.
The trend: AI compute is moving toward tightly coordinated annual platform cycles in which accelerator advances depend on parallel progress in high-bandwidth memory and system integration.
NVIDIA Spectrum-X Roadmap: Spectrum-X800 Ultra in 2025 - 51.2T - 512 Radix - CX8 800 SuperNIC for 100's of thousands of GPUs Spectrum-X1600 in 2026 - 102.4.2T - 512 Radix - CX9 1600G SuperNIC for millions of GPUs [image]
During the keynote Jensen mentioned joules per token pre-Blackwell, 17,000J for GPT-4 Now 0.4J per token with Blackwell Insane efficiency improvement, but time to switch to Tokens per Joule (TPJ) since the number is getting so small, which is 2.5 tok/J
This @NVIDIA keynote is the most painful, meandering keynote I have had to watch since about 2016/2017 GTC keynotes. Some of the sentences are actual complete nonsense. Others are mostly filler. Waiting for anything meaningful.
Nvidia's production Blackwell board. It's 20,000 TFLOPS of FP4. Interested to see what Blackwell RTX 5000 GPUs manage to achieve in PC form factors for AI inference and PC gaming [image]
Watch our CEO Jensen Huang's live keynote at #COMPUTEX2024. Discover how the era of #AI is driving a new industrial revolution across the globe. https://www.youtube.com/...
As everyone knows, Rubin R100 GPUs will use a 4x reticle design made using TSMC CoWoS-L packaging on the N3 process node. TSMC plans for up to 5.5x reticle size chips by 2026 which would feature a 100x100mm substrate and allow for up to 12 HBM sites. https://wccftech.com/...
NVIDIA's Computex 2024 keynote Main Points: -> Announcement of The New Rubin Architecture & -> NVIDIA's Extensive AI-Focused CUDA Libraries -> IT Industry To Evolve Into a $100 Trillion Market -> “The more you spend, the more you save.” It was a delight to watch! [image]
JUST IN: Nvidia $NVDA CEO Jensen Huang said today at the Computex trade show in Taiwan that Nvidia's next-generation AI chip platform will be called Rubin and roll out in 2026 The Rubin family of chips will include a new GPU, CPU, as well as networking chips - Reuters [image]
But, but, but ... what about “The Great Stagnation? ” “NVIDIA's NVLink 6 switch is expected to offer data transfer speeds of up to 3600 GB/s, while the CX9 SuperNIC will provide network speeds up to 1600 Gb/s” https://www.guru3d.com/... [image]