A look at the history of generative AI and developments that paved the way for breakthroughs, including CUDA, convolutional neural networks, and transformers
A new class of incredibly powerful AI models has made recent breakthroughs possible. — Progress in AI systems often feels cyclical. Tweets: @arstechnica Tweets: @arstechnica : We're in the early stages of a revolution that could be as profound as Moore's Law, and what's yet to come is both exciting and terrifying. So why is it all happening now? https://arstechnica.com/... Expand More For Next Unexpand More For Next
Context & Ripple Effects
This retrospective answers the 'why now' question behind the current wave by tracing three converging threads: Nvidia's CUDA GPU computing platform, which made large-scale model training practical; convolutional neural networks, which anchored the earlier deep learning productization wave; and transformer architectures, which moved from language into computer vision and directly enabled today's generative models.
The piece lands at a moment when the field is split between fervor and skepticism — experts have already cautioned that near-term promise may be more modest than tools like ChatGPT suggest — and when analysts are flagging that generative AI's benefits flow through a structurally small set of companies. Framing the breakthrough as the product of stacked infrastructure and architecture choices makes that concentration legible rather than accidental.
First-order effects
- Nvidia's position is reframed from chip vendor to foundational layer: because CUDA enabled the training runs behind recent breakthroughs, its software platform becomes the de facto dependency of the generative AI boom.
- The transformer's migration from language models into computer vision extends the same architecture across modalities, widening who can build on it.
Second-order effects
- Challengers named in prior coverage — Google, Amazon, Graphcore, Cerebras — face a higher bar: displacing Nvidia means competing against CUDA's installed role in the training stack, not just its silicon.
- Because capability depends on scarce compute plus proprietary-scale models, the structural issue of power concentrating among a few companies sharpens as adoption spreads.
Third-order effects
- If the pattern holds, industry structure follows the stack: whoever controls the compute platform and the dominant architectures captures the value of every downstream application, making access to training infrastructure the real competitive moat.
- The same scale that makes generative output cheap per-unit also makes previously small-scale harms practical at massive scale, pushing ethical and legal questions from academic debate toward regulatory agenda.
The trend: Generative AI's trajectory is being set by the co-evolution of hardware platforms and model architectures, with control of that combined stack determining how widely — and how narrowly — its gains are distributed.