An overview of the ML software development industry over the past decade: a decline of Nvidia's CUDA monopoly, PyTorch overtaking Google's TensorFlow, and more
the CUDA monopoly is nowhere close to being broken and CUDA will continue to be the key dependency for PyTorch. As a data point, Triton isn't the first rally — multiple vendors use the XLA compiler as a rally point. https://twitter.com/... @wholemarsblog : How Nvidia's CUDA Monopoly In Machine Learning Is Breaking - OpenAI Triton And PyTorch 2.0 https://www.semianalysis.com/ ... Calum Chace / @cccalum : The idea that ChatGPT spells doom for Google Search is almost certainly wrong. But Google's predominance in advanced AI is undermined by its failure to keep TensorFlow more popular that PyTorch. https://www.semianalysis.com/ ... Tren Griffin / @trengriffin : “Nvidia's FLOPS have increased multiple orders of magnitude by leveraging Moore's Law, but primarily architectural changes such as the tensor core and lower precision floating point formats. In contrast, memory has not followed the same path.” https://www.semianalysis.com/ ... @pommedeterre33 : An under appreciated fact (slightly exaggerated in article below): With current move to @openai Triton, PyTorch 2.0 makes Nvidia software moat a bit weaker. AMD & Intel are working to support Triton, easier than reinventing whole CUDA / Bias stack. https://www.semianalysis.com/ ... @jreuben1 : Breaking Nvidia's CUDA Monopoly In ML https://www.semianalysis.com/ ... - PyTorch 2.0 (Dynamic Shapes) - PrimTorch (~250 primitive operator building blocks) - TorchDynamo (partial / guarded graph capture + JIT recapture: ingest any PyTorch user script and generate an FX graph) Mathieu Orhan / @ai_unleashed : Hard to summarise this amazing post. 1) The memory-wall and how much it takes to correctly use a A100 2) PyTorch is winning and 2.0, along with OpenAI Triton, makes it realistic that actual competition with Nvidia emerges https://www.semianalysis.com/ ... @jreuben1 : - AOT Autograd - TorchInductor compiler (Wrapper Codegen) - OpenAI Triton (LLVM IR direct to PTX without cuBLAS / CudNN) - a good article ! @pommedeterre33 : This is the very reason why we bet on Triton instead of writing custom CUDA code for Kernl library, plus at the end when things do not work you review PTX instructions anyway... The main current advantage of CUDA is certainly maturity, but for how much time... https://twitter.com/...