AWS is launching new EC2 G4 instances with support for Nvidia's T4 GPUs, which optimize for running AI models and use ray tracing tech, in the coming weeks
Frederic Lardinois / TechCrunch :
Context & Ripple Effects
AWS has been assembling a menu of accelerator options for years — FPGA instances for GPU-style workloads in 2016, then Elastic Inference to bolt GPU acceleration onto ordinary EC2 instances in late 2018. The missing piece was a dedicated inference-tuned GPU instance, and Nvidia supplied it when it announced the Tesla T4 as an inference-focused data center chip six months ago.
G4 is where those two threads meet: AWS productizing Nvidia's newest inference silicon as a first-class EC2 SKU rather than an add-on service. It matters because inference is where most deployed AI workloads live, and until now customers had to choose between over-provisioned training-class GPUs or Elastic Inference's attached-accelerator model.
First-order effects
- Customers running trained models get a purpose-built instance instead of paying for training-grade GPUs, directly competing with AWS's own Elastic Inference pitch of cutting deep learning costs by up to 75%.
- Nvidia gains a second major distribution channel for the T4 within months of its launch, putting the chip in front of every EC2 customer without a separate procurement cycle.
Second-order effects
- Elastic Inference's role narrows to bursty or small-model workloads while G4 absorbs steady-state inference traffic, forcing AWS to position the two offerings against each other within its own catalog.
- The G4/P4 cadence — T4 today, and P4 instances with Nvidia's A100 Tensor Core GPUs following in late 2020 — shows each Nvidia generation becoming a new AWS rental tier almost immediately.
Third-order effects
- Cloud GPU pricing stratifies by workload class — training versus inference versus general compute like Arm-based options such as Graviton2-powered C6gn instances — turning accelerator selection into a procurement discipline rather than a single GPU choice.
- Once GPU capacity is rented in fine-grained, workload-specific SKUs, it becomes scarce enough to be priced dynamically — a pattern visible much later when AWS raised Nvidia GPU prices in its EC2 Capacity Blocks service while leaving its own Trainium chips untouched.
The trend: Hyperscalers are converting each of Nvidia's chip generations into ever more granular rental SKUs, splitting AI compute into distinct training and inference markets with their own pricing.