/
Navigation
Chronicles
Browse all articles
Explore
Semantic exploration
Research
Entity momentum
Nexus
Correlations & relationships
Story Arc
Topic evolution
Drift Map
Semantic trajectory animation
Posts
Analysis & commentary
Pulse API
Tech news intelligence API
Browse
Entities
Companies, people, products, technologies
Domains
Browse by publication source
Handles
Browse by social media handle
Detection
Concept Search
Semantic similarity search
High Impact Stories
Top coverage by position
Sentiment Analysis
Positive/negative coverage
Anomaly Detection
Unusual coverage patterns
Analysis
Rivalry Report
Compare two entities head-to-head
Semantic Pivots
Narrative discontinuities
Crisis Response
Event recovery patterns
Connected
Search: /
Command: ⌘K
Embeddings: large
TEXXR

Chronicles

The story behind the story

days · browse · Enter similar · o open

Meta details its two new data center scale clusters, both containing 24,576 Nvidia H100 GPUs that the company is using for AI workloads like training Llama 3

Charlotte Trueman / DatacenterDynamics :

DatacenterDynamics Charlotte Trueman

Context & Ripple Effects

Meta's new clusters extend its earlier AI Research SuperCluster, which was planned around 16,000 GPUs for machine-learning training. The reported build makes infrastructure scale a central part of Meta's effort to train Llama 3.

The move also sits alongside cloud providers' H100-based systems, including Google's A3 GPU supercomputer VM, showing that leading AI developers and cloud platforms are building around the same accelerator generation.

First-order effects

  • Meta gains two dedicated 24,576-H100 clusters for AI workloads, increasing the compute available for training Llama 3 and related models.
  • Nvidia supplies the GPUs at the core of both installations, reinforcing H100's role in large-scale model-training deployments.

Second-order effects

  • Meta's in-house capacity reduces its need to rely solely on external cloud training infrastructure, while raising the bar for peers seeking comparable frontier-model training capability.
  • The deployment strengthens demand for the data-center networking, power, cooling, and systems engineering required to operate GPU clusters at this scale.

Third-order effects

  • If this build-out pattern persists, frontier AI development will become more concentrated among organizations able to finance and operate very large dedicated clusters.
  • The industry is moving toward larger GPU superclusters as a recurring unit of AI infrastructure, making compute access and operational execution more consequential competitive variables.

The trend: AI leaders are treating large, dedicated GPU clusters as strategic infrastructure for training increasingly capable foundation models.

Discussion

  • @documentingmeta @documentingmeta on threads
    “...Meta bought up all those chip orders, so they got into all the GPUs before everyone else did, and they revealed last quarter they had this astronomical fleet because of that specific quarter..."" There's still people who think billions of dollars were going into the “metavers…
  • @firstadopter Tae Kim on threads
    Nvidia's data center stack is not just GPUs: “We also optimized our network routing strategy in combination with NVIDIA Collective Communications Library (NCCL) changes to achieve optimal network utilization.” “features an NVIDIA Quantum2 InfiniBand fabric.” …
  • @stshank Stephen Shankland on threads
    Fascinating (to me anyway) details on how Facebook builds its data centers, including how crditical optimizing everthing is to get consistent high-speed communications.  350,000 Nvidia H100 GPUs today and the equivalent of 600,000 H100s on the way.  Interesting how that compares …
  • @ylecun Yann LeCun on x
    The computing infrastructure for Llama-3 training.
  • @giffmana Lucas Beyer on x
    I read all of this, and my brain just goes: “I really hope they are letting Hugo train some yolo runs on it without too much process headache, given his track record!”
  • @abursuc Andrei Bursuc on x
    Strategies to attract talent for ML at scale: - most companies: hire specialized recruiters, linkedin announcements, ads, etc. - Meta: let us tell you about our new SoTA H100 cluster 🙃
  • @fb_engineering @fb_engineering on x
    Introducing our two new 24k GPU clusters! These clusters will support our current and next-gen AI models, including Llama 3, and help us push the boundaries of AI research. https://engineering.fb.com/...
  • @_akhaliq @_akhaliq on x
    Meta introduces two new 24k GPU clusters These clusters will support current and next-gen AI models, including Llama 3 [image]
  • @ratulm Ratul Mahajan on x
    ML training is a networking problem.