MLCommons shares the results from its MLPerf 4.0 inferencing benchmarks, which added Llama 2 70B and Stable Diffusion XL; PCs with Nvidia GPUs came out on top
no Blackwell submissions yet, sorry Karl Freund / Forbes : Nvidia Sweeps AI Benchmarks While AMD Misses The Boat. Again. Intel : Intel Gaudi 2 Remains Only Benchmarked Alternative to NV H100 for GenAI Performance Cliff Robinson / ServeTheHome : NVIDIA MLPerf Inference v4.0 is Out Hassan Mujtaba / Wccftech : NVIDIA Hopper H200 GPU Continues To Dominate In Latest MLPerf 4.0 Results: Up To 3x Gain In GenAI With TensorRT-LLM Hassan Mujtaba / Wccftech : Intel Gaudi 2 Accelerators Showcase Competitive Performance Per Dollar Against NVIDIA H100 In MLPerf 4.0 GenAI Benchmarks X: @aiatmeta : Announced today: @MLCommons is adopting Meta Llama 2 70B for MLPerf Inference v4.0 ➡️ https://mlcommons.org/... The benchmark is a standard for measuring ML & AI performance across domains and we're excited to support the community in using Llama 2 as part of the benchmark suite. @mlcommons : The @MLPerf Inference v4.0 benchmark suite includes our largest model to date, @Meta's Llama 2 70B large language model with more than 70 billion parameters. Learn more about the selection process, and performance metrics in the benchmark: https://mlcommons.org/... #GenAI @nvidiadc : In the latest #MLPerf benchmarks, NVIDIA H200 Tensor Core GPUs running TensorRT-LLM software delivered the fastest Llama 2 70B inference performance in MLPerf's biggest test of #generativeAI to date. https://blogs.nvidia.com/... @intel : The @MLPerf results are in! We're raising the bar with competitive solutions for your high-performance, high-efficiency deep learning inference needs — even on challenging LLMs. Read more about the results. https://www.intel.com/... #IntelXeon #IntelGaudi #Intel [video] Tony Mongkolsmai / @tonymongkolsmai : @MLPerf results are back baby! Always impressed by my colleagues pushing out performance on the #IntelGaudi 2 AI Accelerators. MLPerf submissions are hard, you have to get it working and make it fast. Two things that aren't trivial when you talk about the scale of things like... @mlcommons : @MLPerf Inference v4.0 results are out! This round includes two new benchmarks focused on gen AI: @Meta's Llama 2 70B model and @StableDiffusion XL. See the complete results and learn more: https://mlcommons.org/... #GenAI #LLM Lauren Wagner / @typewriters : One of the best things I've done all year is collaborate with @MLCommons on AI governance and benchmarking They're my favorite kinds of people to work with: pragmatic, optimistic about the future of technology and peoples' ability to shape it, and focused on building solutions
Context & Ripple Effects
MLPerf's prior LLM-focused results had already put Nvidia's H100 ahead of Intel's Gaudi2 on a smaller model workload. MLPerf 4.0 extends that comparison to larger language-model inference workloads and image generation, making the results more relevant to emerging generative-AI deployments.
The absence of Blackwell entries means this is a snapshot of the available submission set, rather than a comparison of Nvidia's next architecture. It also separates accelerator leadership from the distinct question of performance per dollar, where Intel's Gaudi2 appeared competitive against H100.
First-order effects
- Nvidia's H200/TensorRT-LLM stack posts the fastest submitted Llama 2 70B inference result, while Nvidia-equipped PCs lead the reported PC results.
- Intel gains a benchmarked alternative in Gaudi2 on value-oriented comparisons, but AMD has little visible MLPerf 4.0 evidence to counter Nvidia's results; Blackwell has no submitted result yet.
Second-order effects
- Enterprise buyers evaluating Llama 2 70B or Stable Diffusion XL receive more workload-specific evidence favoring Nvidia hardware and software together, strengthening the practical pull of its platform.
- Intel's competitive performance-per-dollar positioning raises the importance of total deployment cost rather than peak throughput alone, while rivals face pressure to submit comparable results on the same workloads.
Third-order effects
- As benchmarks incorporate production-relevant generative models, accelerator competition is likely to be judged increasingly at the workload-and-software-stack level, not by chip specifications alone.
- If participation remains uneven, public benchmarks may reinforce incumbent platform advantages by shaping buyer shortlists; broader, continuously updated testing could temper that effect, as automated LLM inference tracking illustrates.
The trend: AI accelerator competition is shifting toward workload-specific inference economics, where hardware performance, model support, and software integration jointly determine buyer choices.