Benchmarking AMD's MI300X and Nvidia's H100 and H200; in theory, AMD's GPU has advantages in specs and Total Cost of Ownership, but software bugs hold it back
Intro — SemiAnalysis has been on a five-month long quest to settle the reality of MI300X. In theory, the MI300X …
SemiAnalysis Dylan Patel
Related Coverage
- AMD's poor software optimization is letting Nvidia maintain an iron grip over AI chips TechSpot · Zo Ahmed
- I cannot express my frustration with Intel and AMD at this point. It's 10 years and counting since deep learning should have been on everybody's radar and they still cannot seem to grasp that you need _software support_ to make people use your chips. — This NVidia monopoly on #ai hardware is not good for anybody. … @pbloem@sigmoid.social · Peter Bloem
Analysis
Discussion
-
@halvarflake
@halvarflake
on bluesky
Ok, reading semianalysis.com/2024/12/22/ m... it seems that *most* of AMDs problems could indeed be remedied by having a team of 4-10 very strong engineers focus on always keeping default PyTorch fast on AMD hardware, with nightly regression tests.
-
@cudahandbook
Nicholas Wilt
on x
NVIDIA to churn the hardware instruction set with abandon, sometimes even committing featurecide and relying on the PTX translator to emulate instructions that were removed. The PTX translation code is in the driver as well as the offfline toolchain (ptxas) and is 3/x
-
@cudahandbook
Nicholas Wilt
on x
supercomputers in the world. So yeah, CUDA is a deep, deep moat. I get pretty offended when folks intimate that any luck was involved. We knew exactly what we were doing and why. /fin
-
@cudahandbook
Nicholas Wilt
on x
why I was making sure it ran on Windows as well as Linux. It was healthy for the code base. Today, NVIDIA has parlayed CUDA's Windowa support into a monopoly position in GPU workstations, because 1,200 workstation apps use CUDA. Another pillar is PTX, which enables 2/x
-
@cudahandbook
Nicholas Wilt
on x
multithreaded, so it can exploit modern multicore CPUs for performance gains proportional to the core count. Another triumph of software engineering. All of this great software runs on a span of platforms from tiny SOCs for cars and drones and robots, to the biggest 4/x
-
@cudahandbook
Nicholas Wilt
on x
CUDA's software stack has a few distinct pillars that are triumphs of software engineering (let alone software architecture). The driver API was built in C, portable across both operating systems and CPU architectures. Across the 6 years I worked on CUDA, no one questioned 1/x
-
@realgeorgehotz
George Hotz
on x
“AMD's software experience is riddled with bugs rendering out of the box training with AMD is impossible.” “It's not just that it's immature software, they need to change how they do development.” Good luck @dylan522p!
-
@youjiacheng
You Jiacheng
on x
> Tensorwave, the largest AMD GPU Cloud has given GPU time for free to a team at AMD to fix software issues, which is insane given they paid for the GPUs. This is INSANE...
-
@devnull_0
Indrajit Bhosale
on x
Recently tried to (painstakingly) write a convolution layer kernel using AWS Trainum with their NKI programming model (their latest and greatest), only to realize CUDA moat is well and truly alive!
-
@dylan522p
Dylan Patel
on x
Our 5-month journey conducting independent analysis & benchmarking of AMD MI300X vs Nvidia H100 + H200 Detailed, open source low-level benchmarks performance vs TCO Comprehensive public recommendations It's not just immature software, they need to change how they do development
-
@vintrotweets
@vintrotweets
on x
> Tensorwave, the largest AMD GPU Cloud has given GPU time for free to a team at AMD to fix software issues, which is insane given they paid for the GPUs. glad we got the tinybox green
-
@benthompson
Ben Thompson
on x
This is incredible work. And yeah, AMD has sucked at software for 50 years.
-
@benbajarin
Ben Bajarin
on x
Repeat after me the reason Nvidia GPUs remain so well positioned - massive (and growing installed base) and architectural compatibility. Good analysis below.
-
@theycallmemr_
Saurabh Dash
on x
“AMD” is insane because it appears to be one of the most legitimately dangerous companies with the potential to gigafry the market but exclusively employs literal turbonormies who unironically want to like design x86 processors and basically get oneshotted by their own drivers.
-
@lemire
Daniel Lemire
on x
NVIDIA's GPUs excel due to their mature CUDA software stack, which provides better performance in practical scenarios even when they have inferior hardware specs. via @GeimanThiesen [image]
-
r/nvidia
r
on reddit
MI300X vs H100 vs H200 Benchmark Part 1: Training - CUDA Moat Still Alive