MIT spinoff Liquid AI debuts its non-transformer AI models LFM-1B, LFM-3B, and LFM-40B MoE, claiming they achieve “state-of-the-art performance at every scale”
The release matters because it puts Liquid AI's performance claim alongside deployable model-size choices, rather than positioning the company solely around its underlying neural-network approach.
First-order effects
Liquid AI gains a public product lineup—LFM-1B, LFM-3B, and LFM-40B MoE—against which developers and evaluators can test its non-transformer performance claims.
Prospective users now have models at multiple scales to assess for their own workloads, while Liquid AI must substantiate its “state-of-the-art” positioning through real-world evaluation and adoption.
Second-order effects
Model providers built around transformer architectures face another architectural benchmark point: a credible alternative can shift comparisons toward task performance, deployment fit, and efficiency rather than architecture alone.
Buyers evaluating AI models have more reason to run workload-specific trials instead of treating model family or parameter scale as a sufficient proxy for value—an instance of the push toward smaller, cheaper capable models.
Third-order effects
If non-transformer systems repeatedly deliver competitive results across sizes, the AI-model market could become less architecturally standardized and more open to specialized designs.
That would increase the importance of transparent, application-level evaluation: broad performance claims alone will be less decisive as enterprises compare models against operational requirements.
The trend: This is one data point in AI's shift from a transformer-centered race toward competing architectures optimized for practical capability and deployment economics.
This is the proudest release of my career :) At @LiquidAI_, we're launching three LLMs (1B, 3B, 40B MoE) with SOTA performance, based on a custom architecture. Minimal memory footprint & efficient inference bring long context tasks to edge devices for the first time! [image]
> no weights (despite comparing with open models) > SoTA base LLM perf, claims new Pareto frontier > @MParakhin bullish, on the advisory board > “Human preference optimization techniques have not yet been applied to our models, extensively.” that's one sure way to annoy me, argh …
LFMs are Memory efficient LFMs have a reduced memory footprint compared to transformer architectures. This is particularly true for long inputs, where the KV cache in transformer-based LLMs grows linearly with sequence length. [image]
I've used these models. They are very real. In a game that is changing at increasingly greater frequency, this has the potential to fundamentally change the game.
What Language LFMs are not good at today: Zero-shot code tasks Precise numerical calculations Time-sensitive information Counting r's in the word “Strawberry”! Human preference optimization techniques have not yet been applied to our models, extensively.
First impressions on Liquid, the first non-GPT LLM that is good: * Extremely fast, even when there is a lot of context * Good world fact memorization * Logic is not all the way to o1 (of course), but seems comparable to GPT-4o or Llama 405b https://playground.liquid.ai/
i usually would never retweet these corporate pr releases unless they share some real details, but a long time ago one of their investors pitched their foundation model idea to me and i was privately very skeptical so, publicly, i'll admit that it seems like i was wrong!
new transformers killer in town! Been excited about @LiquidAI_ since I talked to @Plinz in April. Now they've finally launched LFMs! Shots fired: - Better MMLU, ARC, GSM8K than 1B/3B models, comparable-to-slightly-better in the 30-70B weight class. Notably, including Apple's [ima…
Last time I posted about the specific model was the Phi family release - it was a breakthrough. Today - another one, please take note. @SebastienBubeck himself in our private correspondence said: “The Liquid models are strong. Wow.” - sorry, Seb, oversharing, I know
LFM-1B performs well on public benchmarks in the 1B category, making it the new state-of-the-art model at this size. This is the first time a non-GPT architecture significantly outperforms transformer-based models. [image]
LFM-3B delivers incredible performance for its size. It positions itself as first place among 3B parameter transformers, hybrids, and RNN models, but also outperforms the previous generation of 7B and 13B models. It is also on par with Phi-3.5-mini on multiple benchmarks, while […
What Language LFMs are good at today: General and expert knowledge Mathematics and logical reasoning Efficient and effective long-context tasks A primary language of English, with secondary multilingual capabilities in Spanish, French, German, Chinese, Arabic, Japanese, and
LFM-40B offers a new balance between model size and output quality. It leverages 12B activated parameters at use. Its performance is comparable to models larger than itself, while its MoE architecture enables higher throughput and deployment on more cost-effective hardware. [imag…
Congrats to friends @LiquidAI_ , context length scaling is quite something, who needs transformers? Helped with some flops early days, looking forward to seeing the amazing team scale even further 🚀 [image]