Cerebras Systems unveils the world's biggest semiconductor chip that is the size of a large mousepad, with 400K cores, 1.2T transistors, and 18GB of SRAM memory
In the late 1970s, I sat down with technologist … Colm Gorey / Silicon Republic : World's largest semiconductor with 1.2trn transistors could supercharge AI Andy Hock / Cerebras : Introducing Cerebras Systems Cade Metz / New York Times : To Power A.I., Start-Up Creates a Giant Computer Chip Tweets: Soumith Chintala / @soumithchintala : One of the most exciting things you can do with this chip is to design sparse networks, which would be traditionally slow on a GPU / TPU https://twitter.com/... Soumith Chintala / @soumithchintala : the Cerebras chip is a technological marvel — a real, working full-wafer chip with 18GB of register file! It's probably one of the first chips where data “feed” will become the bottleneck, even for fairly modern networks. Congrats @CerebrasSystems ! https://techcrunch.com/... Cerebras Systems / @cerebrassystems : Thanks @DannyCrichton and @TechCrunch for this article on the many engineering challenges we overcame on the way to #waferscale integration-a first in the industry! https://techcrunch.com/... Tren Griffin / @trengriffin : “The ‘Wafer Scale Engine’ is 1.2 trillion transistors (the most ever), 46,225 square millimeters (the largest ever), and includes 18 gigabytes of on-chip memory (the most of any chip on the market today) and 400,000 processing cores.” https://techcrunch.com/... Tren Griffin / @trengriffin : Big is the new small. “At eight and a half inches on each side, it is the biggest computer chip the world has ever seen.” https://fortune.com/... https://twitter.com/... Leonid Boytsov / @srchvrs : This is freaking insane: 400K cores on a single cheap connected via a fast 2D-mesh interconnect with 1 cycle latency for same-core memory. This monstrosity runs on merely 15KWh power. https://fortune.com/... Unfortunately, not benchmarks yet. via https://www.linkedin.com/... Matthew Lynley / @mattlynley : Also a good indicator of how long it takes to get a brand new chip out the door. I remember first hearing about these guys in late 2016 and that was just a whisper. https://twitter.com/... Steve Smith / @spsmith1 : Wafers incur defects when circuits are burned into them, and those areas become unusable...Cerebras had to build in redundant circuits, to route around defects in order to still deliver 400,000 working cores, https://twitter.com/...
Context & Ripple Effects
This is the debut of Cerebras' wafer-scale bet: instead of carving one die out of a wafer, the startup keeps the whole wafer intact as a single chip with 400K cores and all memory on-die as SRAM. The pitch is aimed squarely at the bottleneck GPU-based training hit as models grew — data shuttling between chips, not raw FLOPs.
The related coverage shows the bet compounding rather than fading: within months Cerebras shipped the CS-1 system with Argonne National Lab as its first customer, then iterated through the WSE-2 on 7nm in 2021 and the WSE-3 on TSMC's 5nm in 2024 — each generation scaling transistor counts on the same unusual form factor.
First-order effects
- Cerebras moves from stealth to a product roadmap in one step, giving national labs like Argonne an AI-training machine whose on-chip SRAM removes the inter-chip communication overhead that dominates multi-GPU clusters.
Second-order effects
- Sparse-network designs that run poorly on GPUs and TPUs get a hardware home — PyTorch contributor Soumith Chintala flagged exactly this in reaction to the launch — pressuring accelerator vendors to answer on memory-bandwidth grounds rather than core counts.
Third-order effects
- If wafer-scale keeps iterating the way the corpus shows it did (WSE-2 to WSE-3), 'one chip per wafer' becomes a durable third path in AI silicon alongside GPUs and TPUs — later evidenced by Cerebras powering an OpenAI inference tier and filing for a public listing, though the IPO path has already seen one withdrawal.
The trend: AI compute is diversifying beyond the GPU monoculture, with wafer-scale integration emerging as one sustained architectural response as model sizes outrun conventional chip scaling.