ββ days Β· ββ browse Β· Enter similar Β· o open
Alibaba releases Qwen3-Next, a new model architecture optimized for long-context understanding, large parameter scale, and better computational efficiency
the FUTURE of efficient LLMs is here! πΉ 80B params, but only 3B activated per token β 10x cheaper training, 10x faster inference than Qwen3-32B.(esp. @ 32K+ context!) πΉHybrid Architecture: Gated DeltaNet + Gated Attention β best of speed & [image] Emad / @emostaque : Fast β Cheap β Good β I would estimate this cost < $500k of compute to train & outperforms pretty much any model from last year Lots of interesting tech choices in here, will be very suitable for continuous RL & more Hybrid makes a lot of sense as well @kimmonismus : Holy moly, Qwen is cooking! Qwen-3-Next-90b-A3b is next level efficiency [image] Forums: Hacker News : Qwen3-Next
π Introducing Qwen3-Next-80B-A3B β the FUTURE of efficient LLMs is here! πΉ 80B params, but only 3B activated per token β 10x cheaper training, 10x faster inference than Qwen3-32B.(esp. @ 32K+ context!) πΉHybrid Architecture: Gated DeltaNet + Gated Attention β best of speed & [imag…
Fast β Cheap β Good β I would estimate this cost < $500k of compute to train & outperforms pretty much any model from last year Lots of interesting tech choices in here, will be very suitable for continuous RL & more Hybrid makes a lot of sense as well