NVIDIA announces new Tesla P100 GPU with 15B+ transistors, 16GB of High-Bandwidth Memory for deep learning, manufactured using latest 16nm FinFET process
With massive amounts of computational power … Bob Sherbin / The Official NVIDIA Blog : Live: Jen-Hsun Huang Kicks Off NVIDIA's 2016 GPU Technology Conference Chris Williams / The Register : Inside Nvidia's Pascal-powered Tesla P100: What's the big deal? Brad Chacos / PCWorld : Nvidia's Pascal GPU tech specs revealed: Full CUDA count, clock speeds, and more Jeff Kampman / The Tech Report : Pascal makes its debut on Nvidia's Tesla P100 HPC card Nathan Ingraham / Engadget : NVIDIA's insane DGX-1 is a computer tailor-made for deep learning Devin Coldewey / TechCrunch : NVIDIA announces a supercomputer aimed at deep learning and AI Bruno Ferreira / The Tech Report : Nvidia DGX-1 uses eight Tesla P100s to speed up deep learning Agam Shah / PCWorld : Nvidia's DGX-1 supercomputer packs the horsepower of 250 servers Dean Takahashi / VentureBeat : Nvidia creates a 15B-transistor chip for deep learning Brad Chacos / PCWorld : Nvidia's monstrous Pascal GPU is packed with cutting-edge tech and 15 billion transistors Hilbert Hagedoorn / Guru3D.com : Nvidia announces Tesla P100 data-center GPU Hassan Mujtaba / WCCFtech : NVIDIA 16nm Pascal Based Tesla P100 With GP100 GPU Unveiled …
Context & Ripple Effects
At GTC 2016, Nvidia is positioning the Tesla P100 not as a graphics part but as the opening move of a dedicated AI-compute business: the Pascal debut pairs 15B+ transistors on 16nm FinFET with 16GB of High-Bandwidth Memory aimed squarely at deep learning, and ships alongside the DGX-1, an eight-P100 supercomputer built for the same workload.
That framing set the template for everything in the related coverage — the Titan V later claimed 9x the computing capability of Pascal-based Titan X, the A100 became a ~$10K staple of generative AI work with Nvidia holding an estimated 95% machine-learning GPU share, and the line culminated in rack-scale platforms like the DGX GH200 supercomputer. The P100 is where that arc starts.
First-order effects
- Deep-learning labs get their first HBM-equipped accelerator: 16GB of High-Bandwidth Memory plus the 16nm FinFET node directly addresses the memory-bandwidth wall that neural-network training was hitting on prior-generation parts.
- Nvidia converts the chip launch into a systems sale immediately — the DGX-1 packages eight P100s as a turnkey deep-learning appliance, moving the company up the stack from component vendor to integrated-system seller.
Second-order effects
- Cloud providers and HPC vendors must decide whether to resell Nvidia's boxed systems or build their own P100-based offerings, since the DGX-1 sets a reference price point for what a deep-learning node should cost.
- Rival accelerator makers face a moving target: each Pascal-class generation resets the performance baseline, as seen when Titan V's 110 teraflops was pitched as 9x Pascal just over a year later.
Third-order effects
- If the pattern holds, AI accelerators consolidate into an annual cadence of data-center-first launches tied to one software stack — CUDA — which is how a single vendor reached an estimated 95% share of machine-learning GPUs by the A100 era.
- The chip-to-appliance structure pioneered here points toward AI compute being bought as integrated infrastructure rather than components, ending in rack-scale systems like the DGX GH200.
The trend: Data-center GPUs are evolving from general-purpose chips into annually refreshed, vertically integrated AI systems sold as complete appliances.