/
Navigation
Chronicles
Browse all articles
Explore
Semantic exploration
Research
Entity momentum
Nexus
Correlations & relationships
Story Arc
Topic evolution
Drift Map
Semantic trajectory animation
Posts
Analysis & commentary
Pulse API
Tech news intelligence API
Browse
Entities
Companies, people, products, technologies
Domains
Browse by publication source
Handles
Browse by social media handle
Detection
Concept Search
Semantic similarity search
High Impact Stories
Top coverage by position
Sentiment Analysis
Positive/negative coverage
Anomaly Detection
Unusual coverage patterns
Analysis
Rivalry Report
Compare two entities head-to-head
Semantic Pivots
Narrative discontinuities
Crisis Response
Event recovery patterns
Connected
Search: /
Command: ⌘K
Embeddings: large
TEXXR

Chronicles

The story behind the story

days · browse · Enter similar · o open

Sources: Google is developing a specialized server chip, informally dubbed “Frozen v2”, that integrates its Gemini AI model blueprint into the silicon, for 2028

The Information

Context & Ripple Effects

Google has been extending Gemini from a model program into a broad product layer, following plans to deploy it across much of its product line. Its infrastructure has also advanced through Trillium, the sixth-generation AI chip that powers Gemini 2.0.

Frozen v2 would push that hardware roadmap further by tailoring silicon to a particular model blueprint. It follows reported work on Icefish that splits manufacturing between Samsung for a memory I/O die and TSMC for the compute engine, underscoring how central packaging and memory design have become to AI systems.

First-order effects

  • Google’s chip and Gemini teams would need to coordinate model architecture and hardware design much earlier; the reported 2028 target makes this a long-range infrastructure program rather than a near-term product launch.
  • A chip built around Gemini’s blueprint could give Google a more purpose-built internal platform for workloads that match that model design, while reducing flexibility if the model architecture changes materially before deployment.

Second-order effects

  • The program raises the value of specialized memory, packaging, and foundry capabilities in Google’s TPU supply chain, building on the reported multi-supplier Icefish design.
  • Rival cloud and model providers face added pressure to decide where custom silicon delivers enough workload-specific advantage to justify tighter coupling between their model and hardware roadmaps.

Third-order effects

  • If model-specific chips become repeatable, AI infrastructure competition may shift from buying broadly capable accelerators toward co-optimizing models, compilers, memory systems, and silicon as a single stack.
  • That shift could deepen the divide between companies that operate models at sufficient scale to justify bespoke hardware and those that rely on more general-purpose compute, though the payoff depends on model designs remaining stable long enough to reach production.

The trend: This is part of the move toward inference as strategic infrastructure, where leading AI providers co-design hardware around their own model stacks rather than treating compute as a generic input.

Discussion

  • @kimmonismus @kimmonismus on x
    Google may be preparing to freeze parts of Gemini's architecture directly into silicon. Informally called “Frozen v2,” the chip reportedly targets 6-10× more tokens per watt than Google's newest TPUs. Deployment is planned for as early as 2028. The motivation is immediate: [image…
  • @andrewcurran_ Andrew Curran on x
    Google is developing a new chip named ‘Frozen v2’ specifically designed to run the Gemini family of models more efficiently. The new hardware is expected to arrive in 2028. This was originally reported by The Information. Gemini probably helped design this. [image]
  • @amir Amir Efrati on x
    🔥chip news: Google planning to bake LLM architecture into new AI server chips (non-TPUs) to run Gemini 6-10x more efficiently [image]
  • @pmddomingos Pedro Domingos on x
    It'll be obsolete by the time it's in production.
  • @cxcarroll @cxcarroll on x
    This is very well-suited for latency-sensitive applications like voice. While it appears that frontier models may become somewhat commoditized, producing silicon specific to models is going to be reserved for just a couple of the big players.
  • @scottw_grizzle Scott Willis on x
    Interesting solution to the compute problem.
  • @kawzinvests @kawzinvests on x
    Taalas has a public demo called Chat Jimmy running their hardwired Llama 3.1 8B at around 17,000 tokens per second. It's honestly insane. You hit enter and the full response is just there. About 0.03 seconds. It streams faster than you can physically read.
  • @zephyr_z9 @zephyr_z9 on x
    So, this chip makes a lot of sense for real-time voice models and stuff that requires extremely low latency
  • @jukan05 Jukan on x
    So Google is essentially developing Taalas-like chips that bake the model weights directly into the silicon? [image]
  • @stocksavvyshay Shay Boloor on x
    $GOOGL is developing a new AI chip called “Frozen v2” that could run Gemini models nearly 10x more efficiently than its latest TPUs. The chip is targeted for 2028 and would hardwire parts of Gemini to improve speed and efficiency while easing Google's compute constraints. [image]