Nvidia unveils the Grace CPU Superchip, an ARM-based discrete data center CPU with 144 high-performance cores and 1TB of secondary memory, coming in 2023
Nvidia offered details on its Grace central processing unit (CPU) “Superchip” during CEO Jensen Huang's keynote speech at its virtual Nvidia GTC 2022 event.
VentureBeatDean Takahashi
Context & Ripple Effects
Grace began as an Arm server-CPU effort for large-scale neural-network workloads; this announcement turns that earlier Grace server-CPU plan into a more defined discrete data-center product. It broadens Nvidia's role in the server beyond supplying accelerators.
The product also established the CPU component Nvidia later used in the GH200 Grace Hopper pairing and, later, the GB200 configuration that combines Grace with B200 GPUs. The arc is toward systems designed as CPU-and-accelerator units rather than separately selected parts.
First-order effects
Nvidia gains a discrete Arm CPU offering for data-center customers, extending its product scope from GPU acceleration into the host-processing layer.
Grace gives Nvidia a specified 144-core, 1TB-memory platform around which to position its own data-center products scheduled for 2023.
Second-order effects
Nvidia can make the CPU a tighter companion to its accelerators, a direction reflected in the later GH200 Grace Hopper superchip rather than treating the host processor as an external component.
Data-center buyers evaluating Nvidia accelerators gain a CPU option from the same supplier, increasing the appeal of integrated system roadmaps.
Third-order effects
If Nvidia continues carrying Grace into combined products, server competition shifts toward heterogeneous compute packages whose CPU, GPU and memory choices are designed together.
The later Grace Hopper and Grace Blackwell products suggest that the discrete CPU is becoming a building block in Nvidia's broader integrated data-center stack, not a one-off server chip.
The trend: Nvidia is moving from accelerator supplier toward an integrated heterogeneous-compute provider built around paired CPUs, GPUs and memory.
The CPU, as defined by.... Intel: Center of the universe, the answer for all things. Nvidia: A thing that gets your data to my GPU. https://twitter.com/...
It's interesting to hear Jensen talk about the Grace CPU. He's much more focused on how it moves data - especially the memory bandwidth - than focus on single thread performance (which he also says will be very good).
Very interesting if the marketing is close to reality: NVIDIA has developed a variant of the tensor core for speeding up transformers, the Transformer Engine. They swap between FP16 and FP8 on a layer by layer basis. A large cluster of GPUs can train a transformer up to 9x faster…
The Hopper GH100 whitepaper is also out. Still digging through it, but the basics are 144 SMs, with 128 FP32 CUDA cores per SM (versus 64 in GA100). NV is thinking they'll hit clockspeeds around 1.9GHz https://twitter.com/...
NVIDIA has just announced their next-generation GPU architecture: Hopper. The first part out of the gate, the H100 server accelerator, is comprised of 80B transistors, and along with trebling tensor throughput, will draw an eye-popping 700 Watts of power https://www.anandtech.com…
Announcing the #NVIDIAHopper architecture - the next generation of accelerated computing that securely scales diverse workloads in every data center. #GTC22 https://www.youtube.com/... https://twitter.com/...
Nvidia has just announced its new Hopper architecture. The first chip is the H100, designed for big AI performance boosts. It's the first GPU to support PCle Gen5 and utilize HBM3 https://www.theverge.com/... https://twitter.com/...
Nvidia's strategy to displace x86 from the datacenter is finally coming into focus. Nvidia's Grace CPU tied closely to their newest Hopper arch looks like a beast and will cover a wide range of data center workloads. https://t.co/M7m82i6m9Q