Amazon VP of EC2 Dave Brown says AWS is considering using AMD's new MI300 chips and declined to use Nvidia's DGX Cloud chips; Oracle is Nvidia's first partner
Amazon Web Services (AMZN.O), the world's largest cloud computing provider, is considering using new artificial intelligence chips …
ReutersStephen Nellis
Context & Ripple Effects
AWS's position sits at an early point in its shift toward a broader accelerator portfolio rather than a single supplier path. Later coverage shows AWS planning H200 availability while peers including Microsoft and Oracle backed AMD's MI300X deployments.
The immediate importance is commercial as well as technical: AWS was willing to assess AMD hardware but not adopt Nvidia's DGX Cloud service model, while Oracle became Nvidia's initial partner for that offering.
First-order effects
AWS gains an additional AI-chip option to evaluate for its cloud fleet, potentially giving AMD a route into the largest cloud provider's infrastructure decisions.
Nvidia's DGX Cloud launch proceeds without AWS as a partner, concentrating its initial cloud-service relationship with Oracle.
Second-order effects
AMD's prospective AWS evaluation increases pressure on Nvidia to compete not only on accelerators but on the cloud-provider terms and deployment models attached to them.
Cloud customers face a more fragmented set of AI-compute choices as providers can combine supplier chips with their own infrastructure and service layers.
Third-order effects
If hyperscalers continue to mix accelerator suppliers and resist vendor-operated cloud layers, AI infrastructure is likely to be organized around cloud providers' own platforms rather than a single chipmaker's end-to-end service.
The pattern points to compute procurement becoming strategic leverage: suppliers need both competitive silicon and an operating model that fits each cloud provider's control of customer relationships.
The trend: This is an early data point in hyperscalers building multi-vendor AI compute portfolios while preserving control over the cloud service layer.
Do you use Instagram, Facebook or WhatsApp? @Meta VP @AlexisBjorlin shared how all your favorite apps are powered by hundreds of thousands of AMD servers, with the next generation of #EPYC beginning deployment this year. So go on and share that meme. #EPYC has you covered. [image…
Here it is: fourth-gen EPYC Bergamo for cloud-native applications. Up to 128 Zen 4c cores, each of which are 35% smaller than Zen 4 cores. Aimed at leadership performance-per-watt versus the performance-per-core focus of general-purpose EPYC Genoa. [image]
I'm the only one that has brought their Genoa delidded CPU with them. Oh well :) Here's @AMD #Bergamo. Left: 96 core (12*8) Genoa Right: 128-core (16*8) Bergamo. Same IO die. Same uArch, resized L3. Same power. Same socket. Same DRAM Support. Higher Efficiency. [image]
Interesting discussion with @AMD's @LisaSu and @Meta regarding their planned use of Bergamo for #GenerativeAI apps and early tests showing a 2.5x improvement over Genoa. [image]
AMD revealed the GPU-only Instinct MI300X, which has 192GB of HBM3 memory, beating the 80GB capacity of Nvidia's H100. This will allow users to run large language models on fewer GPUs, giving the MI300X leadership total cost of ownership, Su said. [image]
Brad McCredie, AMD's vice president of data center and accelerated processing, explains the building blocks of the Instinct MI300X GPU in response to a question from @CDemerjian. [video]
While the chiplet-based design of the MI300 always left the door open to further CPU/GPU mix & matching, I wasn't sure if AMD was going to do it. But sure enough, they are, with a GPU-only MI300X. The 8 HBM stacks is going to be a big boon for the LLM/GenAI market they're after h…
And finally, AMD will be producing a super-sized, GPU-only version of the MI300 accelerator, the MI300X. Using 24GB HBM3 stacks and 8 CDNA 3 GPU chiplets, the flagship chip will offer 192GB of local memory, ideal for AI/LLM training and other AI workloads https://www.anandtech.co…