Profile of Cerebras, which made the world's largest chip by using a “wafer-scale” approach that offers one possibility for AI chips to keep up with Moore's law
In the race to accelerate A.I., the Silicon Valley company Cerebras has landed on an unusual strategy: go big.
Context & Ripple Effects
Cerebras has spent four years turning an exotic bet into a product line: the Los Altos startup profiled in 2017 as a little-known maker of AI-specific chips unveiled a mousepad-sized semiconductor with 400K cores and 1.2T transistors in 2019, shipped it as the CS-1 system running basic research at Argonne National Lab that November, and by 2022 claimed the largest NLP model trained on a single device, up to 20B parameters.
The New Yorker profile frames what all of that adds up to: rather than shrink transistors, Cerebras keeps Moore's-law-style gains alive for AI workloads by scaling sideways — one giant wafer-scale die instead of many small chips stitched together.
First-order effects
- AI researchers and national-lab buyers gain a single-device alternative to GPU clusters for large-model training, with Argonne already operating the first CS-1 systems.
- Cerebras converts its $112M early funding into differentiated silicon inventory, making it a named supplier of AI training hardware rather than a stealth bet.
Second-order effects
- GPU-based incumbents face pressure to match per-device memory and core counts for training workloads, since wafer-scale SRAM removes the interconnect bottleneck their multi-chip systems must engineer around.
- A proven single-device training path pulls demand toward specialized AI silicon generally, widening the market for AI-specific chip startups beyond the general-purpose accelerator duopoly.
Third-order effects
- If wafer-scale economics hold, AI hardware stratifies between scaling-down transistor roadmaps and scaling-up die area — with customers choosing per workload instead of defaulting to GPUs, and later signals in the corpus (an OpenAI-powered API tier, an AMD server-rack collaboration, quarterly revenue up 94% YoY) suggesting the approach found commercial traction beyond research labs.
- As Moore's law slows, architectural bets like this shift competition among chipmakers from fabrication leadership toward packaging and system-level design, where new entrants can differentiate without owning leading-edge fabs.
The trend: AI compute is diversifying past Moore's-law scaling into radical architectures — wafer-scale dies among them — as training demand outruns what shrinking transistors alone can deliver.