Toronto-based chip startup Taalas, which hardwires AI models into custom silicon to achieve faster inference, raised $169M, bringing its total funding to $219M
Toronto-based chip startup Taalas said on Thursday it had raised $169 million and has developed a chip capable …
ReutersMax A. Cherney
Context & Ripple Effects
Taalas previously emerged from stealth with $50 million raised across two rounds and a plan to unveil an LLM chip. This financing marks a materially larger capital base behind that custom-silicon approach.
The company enters a growing inference-chip cohort: d-Matrix raised funding for inference-optimized chips in 2023, while Taalas was launched by Tenstorrent founder Ljubisa Bajic. The immediate question is whether model-specific hardware can translate technical speed claims into deployable products.
First-order effects
Taalas gains $169 million in new funding, taking total capital to $219 million and extending its capacity to develop its hardwired-model chip approach.
The raise gives Taalas greater visibility with prospective AI-infrastructure buyers and partners seeking faster inference hardware.
Second-order effects
Inference-chip rivals face added pressure to show that their architectures can deliver practical performance advantages, not just secure financing.
Customers evaluating AI-serving infrastructure gain another specialized-hardware option, increasing the importance of benchmarking inference speed and deployment fit across vendors.
Third-order effects
If model-specific silicon proves deployable at scale, AI hardware could split further between broadly programmable accelerators and specialized inference systems.
That outcome would make capital, chip-design expertise, and access to production and deployment partners more decisive barriers for smaller AI-chip entrants.
The trend: This is one data point in the AI hardware strategy split, as startups seek to capture inference demand with architectures tailored more narrowly than general-purpose AI accelerators.
This is an exciting effort. @cerebras has proven that speed matters. Ljubisa and the team have taken the same approach: develop new engineering to make a model go hyper fast. From a manufacturing perspective, @taalas_inc seems like it's tradeoffs are more digestible & scalable
Today, @taalas_inc is unveiling breakthrough inference chips to make AI cheap, fast, and ubiquitous. Read more from our partner, @kvamme. https://quiet.com/...
@sallywf @taalas_inc And in exchange for those 30 (incremental) tape-outs (that takes only 8 weeks), you're running your frontier model 60x faster and 2x+ cheaper. And that's just the 1st generation! The ROI becomes positive very quickly when running inference at scale.
The biggest bottleneck in AI just got obliterated by @taalas_inc. 17,000 tokens/sec. 50x faster than Nvidia's best GPU, at a fraction of the cost and power. Cheap, instant, intelligent AI for everyone is no longer theoretical. This is Taalas' generation one. Their opening act.
AI chip startup Taalas @taalas_inc is showing off a chip that can do 16,000 tps/user on Llama3.1-8B, many multiples of its nearest competitor. The catch? The chip ONLY runs Llama3.1-8B, and a model like DeepSeekR1-671B would need 30 separate tapeouts: https://www.eetimes.com/...
24 dedicated people. $30M spent on development. Extreme specialization, speed, and power efficiency. Today we launch Taalas' first product. Check it out: Details: https://taalas.com/... Demo chatbot: https://chatjimmy.ai/ API: https://taalas.com/...