Meta details its two new data center scale clusters, both containing 24,576 Nvidia H100 GPUs that the company is using for AI workloads like training Llama 3
Meta's new clusters extend its earlier AI Research SuperCluster, which was planned around 16,000 GPUs for machine-learning training. The reported build makes infrastructure scale a central part of Meta's effort to train Llama 3.
The move also sits alongside cloud providers' H100-based systems, including Google's A3 GPU supercomputer VM, showing that leading AI developers and cloud platforms are building around the same accelerator generation.
First-order effects
Meta gains two dedicated 24,576-H100 clusters for AI workloads, increasing the compute available for training Llama 3 and related models.
Nvidia supplies the GPUs at the core of both installations, reinforcing H100's role in large-scale model-training deployments.
Second-order effects
Meta's in-house capacity reduces its need to rely solely on external cloud training infrastructure, while raising the bar for peers seeking comparable frontier-model training capability.
The deployment strengthens demand for the data-center networking, power, cooling, and systems engineering required to operate GPU clusters at this scale.
Third-order effects
If this build-out pattern persists, frontier AI development will become more concentrated among organizations able to finance and operate very large dedicated clusters.
The industry is moving toward larger GPU superclusters as a recurring unit of AI infrastructure, making compute access and operational execution more consequential competitive variables.
The trend: AI leaders are treating large, dedicated GPU clusters as strategic infrastructure for training increasingly capable foundation models.
“...Meta bought up all those chip orders, so they got into all the GPUs before everyone else did, and they revealed last quarter they had this astronomical fleet because of that specific quarter..."" There's still people who think billions of dollars were going into the “metavers…
Nvidia's data center stack is not just GPUs: “We also optimized our network routing strategy in combination with NVIDIA Collective Communications Library (NCCL) changes to achieve optimal network utilization.” “features an NVIDIA Quantum2 InfiniBand fabric.” …
Fascinating (to me anyway) details on how Facebook builds its data centers, including how crditical optimizing everthing is to get consistent high-speed communications. 350,000 Nvidia H100 GPUs today and the equivalent of 600,000 H100s on the way. Interesting how that compares …
I read all of this, and my brain just goes: “I really hope they are letting Hugo train some yolo runs on it without too much process headache, given his track record!”
Strategies to attract talent for ML at scale: - most companies: hire specialized recruiters, linkedin announcements, ads, etc. - Meta: let us tell you about our new SoTA H100 cluster 🙃
Introducing our two new 24k GPU clusters! These clusters will support our current and next-gen AI models, including Llama 3, and help us push the boundaries of AI research. https://engineering.fb.com/...