Alibaba releases the open-weight Qwen3.5 Small Model Series in 0.8B, 2B, 4B, and 9B sizes, claiming the 9B model rivals OpenAI's gpt-oss-120b on some benchmarks
VentureBeatCarl Franzen
Context & Ripple Effects
Alibaba has been building the Qwen line through open-weight releases, from its earlier Qwen3 hybrid reasoning family to a 397B-parameter Qwen3.5 multimodal model. The Small series extends that arc with substantially smaller size options rather than a single flagship release.
The release also sits alongside Alibaba’s effort to argue that model efficiency can narrow the gap with much larger systems, including its Qwen3.5 launch with cost and workload-performance claims. The 9B comparison is a vendor claim limited to some benchmarks, not a general equivalence finding.
First-order effects
Developers and enterprises evaluating open-weight models gain four Qwen3.5 size tiers, creating more deployment options where a smaller model may fit operational constraints better than a flagship-scale release.
Alibaba directly positions the 9B variant against OpenAI’s gpt-oss-120b on selected benchmarks, putting efficiency as well as raw parameter scale at the center of model comparisons.
Second-order effects
Open-weight model buyers can use the new size ladder as additional leverage when comparing capability, operating requirements, and vendor dependence across model options.
Competing open-weight providers will face pressure to substantiate small-model performance claims with clearer task-specific evaluations, since benchmark parity claims alone do not establish broad substitutability.
Third-order effects
If small open-weight models repeatedly deliver useful performance relative to much larger alternatives, model selection may shift further from headline scale toward workload fit, evaluation quality, and integration support.
This strengthens an open-weight ecosystem in which value can increasingly accrue to deployment, tooling, and application-layer complements rather than solely to the largest base model.
The trend: The release is one point in the shift toward efficient open-weight model portfolios that compete on deployability and workload-specific performance, not parameter counts alone.
🚀 Introducing the Qwen 3.5 Small Model Series Qwen3.5-0.8B · Qwen3.5-2B · Qwen3.5-4B · Qwen3.5-9B ✨ More intelligence, less compute. These small models are built on the same Qwen3.5 foundation — native multimodal, improved architecture, scaled RL: • 0.8B / 2B → tiny, fast, [image…
The Qwen 3.5 small model series is now available ollama run qwen3.5:9b ollama run qwen3.5:4b ollama run qwen3.5:2b ollama run qwen3.5:0.8b All models support native tool calling, thinking, and multimodal capabilities in Ollama.
china just mogged @AlexFinn they waited for him to spend his life savings on 317 mac studios with a combined 4.7 petabytes of vram and then they released frontier models that are small enough to run on a single potato
How is this even possible?! Qwen has released 4 new models and the 4B version is almost as capable as the previous 80B A3B one 🤯 And the 9B is as good as GPT OSS 120B while being 13x smaller! - They can run on any laptop - 0.8B and 2B for your phone - Offline and open source
Do you understand what this means? Are you aware how much the world just changed? You can now run frontier intelligence on a potato Your $600 Mac Mini can now run unlimited super intelligence for free. No authoritarian AI companies can cut you off Do this immediately, no
The new Qwen 3.5 by @Alibaba_Qwen running on-device on iPhone 17 Pro. Qwen 3.5 beats models 4 times its size, has strong visual understanding, and can toggle reasoning on or off. The 2B 6-bit model here is running with MLX optimized for Apple Silicon. [video]
Qwen 3.5 Small models - fully open source - beats models 4x it's size - 9B model performs on par with GPT OSS 120B while being 13x smaller - outperforms Gemini 3 flash and Claude sonnet 4.5 on select benchmarks - runs on any laptop - even works on a phone - completely free. [vide…
INCREDIBLE qwen just dropped 4 new qwen3.5 small models: 0.8b, 2b, 4b, and 9b. looking at the benchmarks, they're now matching gemini 3 flash and sonnet 4/4.5 in about half the tests, even on vision this is a big deal. because you can now run them locally, and they're good [image…