MLCommons' MLPerf benchmark, based on a 6B-parameter LLM that summarizes CNN articles: Nvidia's H100s perform best, followed by Intel's Gaudi2 at ~10% slower
An artificial intelligence benchmark group called MLCommons unveiled the results on Monday of new tests that determine how quickly top-of-the-line hardware can run AI models.
Context & Ripple Effects
This result establishes an early large-language-model inference comparison between Nvidia and Intel on a common workload. Later MLPerf inference rounds broadened the models under test to include Llama 2 70B and Stable Diffusion XL, making the benchmark series more representative of varied deployment workloads.
The initial H100–Gaudi2 gap also provides a baseline for Intel's subsequent accelerator push: Gaudi 3 was positioned with higher claimed training and inference performance than H100. Subsequent MLPerf training results continued to place H100 first across the reported tests.
First-order effects
- Nvidia gains an independent benchmark validation for H100 on the tested 6B-parameter summarization workload, strengthening its performance case with infrastructure buyers.
- Intel can point to Gaudi2 as a close alternative in this specific test, but the reported result leaves it behind Nvidia on the benchmark's headline measure.
Second-order effects
- Enterprise buyers and cloud operators gain a comparable data point for accelerator selection, while vendors face pressure to submit and optimize hardware for shared benchmarks rather than rely solely on vendor claims.
- A narrow benchmark gap shifts attention toward workload fit, software support, availability and total operating cost—areas that can determine purchasing even when raw performance differs.
Third-order effects
- As MLPerf expands to larger language and generative-media models, accelerator competition is likely to be judged across a portfolio of workloads rather than by a single peak-performance result.
- The broader direction is toward continuous LLM inference tracking and more workload-specific comparisons, which could make software optimization and cost per useful output as consequential as chip speed.
The trend: AI accelerator competition is moving from headline hardware throughput toward repeatable, workload-specific benchmarking across increasingly diverse models.