In Q2 2026, server-led eSSDs reached 48% of NAND flash shipments as AI workloads shifted from training to inference, driving 5x YoY industry revenue growth
- AI inference workloads drove enterprise SSDs to 48% of global NAND shipments in Q2 2026, nearly double the 26% share a year earlier.
Context & Ripple Effects
Related coverage had already tied AI inference to a projected acceleration in CPU demand, with AMD's inference-led CPU growth outlook marking the compute side of the shift. The latest NAND data shows that the workload transition is now reaching server storage as well.
The broader chip market had been led by advanced chips growing faster than total sales in 2025; that advanced-chip growth now has a clear adjacent demand signal in enterprise SSDs.
First-order effects
- Enterprise SSD suppliers now account for 48% of NAND shipments, versus 26% a year earlier, making server-oriented storage the dominant source of shipment mix change.
- NAND producers receive a sharply stronger revenue signal from inference-oriented server demand, with industry revenue reported up fivefold year over year.
Second-order effects
- Server builders and AI infrastructure buyers must treat enterprise SSD capacity as a core inference-system input rather than a secondary storage choice, expanding the storage portion of deployment budgets.
- AMD and other compute vendors pursuing inference demand gain evidence that workload growth is transmitting beyond processors into the surrounding server component stack.
Third-order effects
- If inference remains the principal AI workload growth driver, the AI infrastructure market will increasingly be defined by a coordinated compute-and-storage input stack rather than accelerator sales alone.
- The shift raises the strategic value of suppliers that can serve persistent, server-scale inference workloads, extending the AI infrastructure supercycle into NAND mix and revenue.
The trend: AI inference is broadening the infrastructure supercycle from high-end compute into the server storage layers required to run workloads at scale.