AMD acquires Toronto-based Taalas, which integrates model weights directly into silicon with the promise to boost inference performance, for an undisclosed sum
Early tech demos show model-specific integrated circuits churning out up to 17,000 tokens a second
The RegisterTobias Mann
Context & Ripple Effects
Taalas arrived at AMD after a $169 million funding round built around hardwiring AI models into custom silicon. The deal extends AMD’s inference push beyond its earlier acquisition of Untether AI’s employee team and software-optimization startup Brium.
AMD has also added model expertise through its Silo AI acquisition, making Taalas a complementary bet on tying models more tightly to the underlying compute rather than treating the chip as a general-purpose accelerator.
First-order effects
AMD gains Taalas’s model-specific silicon approach, adding an inference option designed around integrating model weights directly into hardware.
Taalas moves from a venture-backed chip startup to an AMD asset, giving its technology access to AMD’s broader AI hardware and software organization.
Second-order effects
AMD’s inference offering becomes more dependent on coordinating model design, software optimization, and silicon, raising the value of its prior Silo AI and Brium capabilities alongside accelerator hardware.
Rival AI-chip suppliers face a sharper comparison on inference specialization: general-purpose accelerators must compete not only on raw throughput but on how closely hardware can be tailored to deployed models.
Third-order effects
If AMD can productize Taalas’s approach, AI inference competition shifts further toward vertically integrated stacks in which model architecture, compilers, and silicon are designed together.
The acquisition reinforces a market structure where specialized inference technology is more likely to be absorbed by platform-scale chip vendors than developed independently through commercialization.
The trend: AI inference is moving from general-purpose acceleration toward integrated stacks that co-design models, software, and specialized silicon.
We are pleased to share that Taalas has agreed to join AMD. We built Taalas to rethink AI inference from the ground up: hardware designed around the model, rather than the other way around. The result is the world's fastest and most cost-effective inference silicon. Joining AMD
Taalas buried the lede for the amazing demo of their first tape out. 15k tokens per second with a llama 8b model, ability to scale that up etched onto silicon As models satisfice etching makes sense, particularly ternary.. Try it out https://chatjimmy.ai/ Bullish for $AMD
Excited to announce our agreement to acquire @taalas_inc. Phenomenal team working at the bleeding edge of AI inference. Looking forward to all we will do together. [image]
Huge congrats to Taalas on joining AMD!! @taalas_inc built silicon around the model to make AI inference dramatically faster, cheaper, and more efficient. Bringing that approach inside AMD will be a massive unlock for the next era of AI compute. Proud that @QuietCapital has
So AMD dropped this low-key acquisition! They now own the world's fastest inferencing champion. Taalas has Llama 3.1 8B running @ 17,000 tps!!!! All in hardware! [image]
I did not see that coming. I've argued that AMD needed to buy one of these companies, and quick. Taalas has hard coded models. Interesting to see where the use cases for that lie.
WOW. AMD acquires Taalas. Has to be a high-interactivity play, but it's clearly different than Helios + Cerebras. Taalas is not programmable, so Helios for prefill + Taalas for decode would be pinned to a particular model. That would be an interesting “semi-custom” inference
One of the great privileges of my career has been working with Ljubisa and the Taalas team from the beginning. They took an extraordinarily ambitious idea and made it real in record time. I can't wait to see what they do with AMD behind them. More soon!
Maybe 5,000tps Qwen3.8-27b personal AMD hardware box is coming soon. Taalas runs Llama 3.1 8b with 17,000tok/s It etches the model weight directly on ASIC. Can't wait to see this magic with Qwen or Gemini.