Sources: Google is in talks with Marvell Technology to develop a memory processing unit that works alongside TPUs, and a new TPU for running AI models
Google is in talks with Marvell Technology to develop two new chips aimed at running AI models more efficiently, according to two people with direct knowledge of the discussions.
Context & Ripple Effects
Google’s TPU effort began as an internal machine-learning chip program and has since expanded into successive training and inference-oriented generations. Related coverage says Google has also started pitching TPUs to external customers, raising the strategic value of improving inference efficiency.
The reported Marvell discussions center on adding a memory-focused processor alongside TPUs and a new inference TPU. Subsequent coverage of Icefish describes a split manufacturing model, with Samsung discussed for a memory input-output die and TSMC for the compute engine.
First-order effects
- Google could broaden its TPU design from a principally compute-centric accelerator into a more disaggregated system that explicitly pairs compute with memory-oriented hardware for model serving.
- Marvell would become a prospective design partner in Google’s AI silicon program, while TPU customers would ultimately have another inference-focused hardware configuration to evaluate if talks produce products.
Second-order effects
- A specialized memory component would increase the importance of memory bandwidth, packaging, and input-output design in TPU roadmaps, extending supplier opportunity beyond the main compute die.
- Google’s external TPU push would gain a more tailored inference proposition, increasing pressure on other AI-accelerator platforms to differentiate their own compute-and-memory architectures rather than compete on accelerator performance alone.
Third-order effects
- If this approach is carried into future TPU generations, AI infrastructure may move further toward heterogeneous, multi-chip systems in which memory and compute are separately optimized and sourced.
- That shift could redistribute AI infrastructure value toward memory-interface and integration specialists, while making supply-chain coordination across chip designers and manufacturers a more central competitive capability.
The trend: This is one data point in the move from monolithic AI accelerators toward heterogeneous systems designed around the memory and serving constraints of AI inference.