Alibaba releases its open-weight Qwen3.5 Small Model Series in 0.8B, 2B, 4B, and 9B sizes, claiming the 9B model rivals OpenAI's gpt-oss-120b on some benchmarks
Earlier today, e-commerce giant Alibaba's Qwen Team of AI researchers, focused primarily on developing and releasing to the world …
VentureBeatCarl Franzen
Context & Ripple Effects
Alibaba is extending the Qwen3.5 line from its earlier 397B-parameter open-weight multimodal release down to models sized for materially lighter deployment. That makes the small-model series a product-line expansion rather than a standalone benchmark announcement.
The move continues Alibaba's Qwen strategy of publishing open-weight models across architectures and scales, following its open-weight hybrid reasoning-model family. The salient claim is not that small models replace frontier systems universally, but that a 9B option can be competitive on some tests.
First-order effects
Developers can evaluate and deploy Qwen3.5 variants from 0.8B to 9B weights, giving teams more choices where hardware limits, latency, or operating cost rule out much larger models.
Alibaba gains a concrete efficiency narrative for Qwen: its 9B model's claimed results against OpenAI's gpt-oss-120b put benchmark attention on capability per parameter, subject to independent validation and workload-specific testing.
Second-order effects
Open-weight model providers and hosted-model vendors face greater pressure to show that larger models deliver enough incremental quality to justify their compute and serving requirements.
Enterprise buyers can use smaller credible alternatives to test task-specific deployments locally or with different hosting partners, strengthening their leverage in model selection and pricing discussions.
Third-order effects
If small open-weight models repeatedly meet practical quality thresholds, differentiation shifts away from raw parameter scale toward tooling, deployment controls, fine-tuning, and distribution.
The Qwen roadmap suggests a bifurcated model market: frontier-scale releases for broad capability alongside compact weights for economical deployment, with benchmark claims increasingly requiring task-level validation to determine purchasing value.
The trend: This is part of the AI industrialization trend in which model vendors compete on deployable capability per unit of compute, not only headline model size.
🚀 Introducing the Qwen 3.5 Small Model Series Qwen3.5-0.8B · Qwen3.5-2B · Qwen3.5-4B · Qwen3.5-9B ✨ More intelligence, less compute. These small models are built on the same Qwen3.5 foundation — native multimodal, improved architecture, scaled RL: • 0.8B / 2B → tiny, fast, [image…
How is this even possible?! Qwen has released 4 new models and the 4B version is almost as capable as the previous 80B A3B one 🤯 And the 9B is as good as GPT OSS 120B while being 13x smaller! - They can run on any laptop - 0.8B and 2B for your phone - Offline and open source
The new Qwen 3.5 by @Alibaba_Qwen running on-device on iPhone 17 Pro. Qwen 3.5 beats models 4 times its size, has strong visual understanding, and can toggle reasoning on or off. The 2B 6-bit model here is running with MLX optimized for Apple Silicon. [video]
The Qwen 3.5 small model series is now available ollama run qwen3.5:9b ollama run qwen3.5:4b ollama run qwen3.5:2b ollama run qwen3.5:0.8b All models support native tool calling, thinking, and multimodal capabilities in Ollama.
The new Qwen on device models receiving praise from Elon These are densely intelligent and replace the need for APIs for general tasks Within a year these could be Opus level as further distillation and training techniques improve They can even be embedded in web apps [image]
Big leap for on-device AI. Here Qwen 3.5 2B (6-bit) running on iPhone 17 Pro, MLX-optimized. Outperforms models 4X its size with strong visual intelligence. Powerful AI, now truly mobile. Opportunity for so many new products. [video]
Alibaba Qwen 3.5 AI models have become unbelievable for their size! Soon we will enter the age of intelligent devices, when eventually an Apple Watch will have the expertise of a PhD on any topic! Then imagine running your AI agents on Raspberry Pis with Qwen 3.5! Not far away!
Do you understand what this means? Are you aware how much the world just changed? You can now run frontier intelligence on a potato Your $600 Mac Mini can now run unlimited super intelligence for free. No authoritarian AI companies can cut you off Do this immediately, no
Qwen 3.5 Small models - fully open source - beats models 4x it's size - 9B model performs on par with GPT OSS 120B while being 13x smaller - outperforms Gemini 3 flash and Claude sonnet 4.5 on select benchmarks - runs on any laptop - even works on a phone - completely free. [vide…
INCREDIBLE qwen just dropped 4 new qwen3.5 small models: 0.8b, 2b, 4b, and 9b. looking at the benchmarks, they're now matching gemini 3 flash and sonnet 4/4.5 in about half the tests, even on vision this is a big deal. because you can now run them locally, and they're good [image…
china just mogged @AlexFinn they waited for him to spend his life savings on 317 mac studios with a combined 4.7 petabytes of vram and then they released frontier models that are small enough to run on a single potato
Qwen 3.5 small models (0.8B-9B) with base models released 🫡 A 4B multimodal model that runs on a single 4090 — this is what makes LLM research accessible to PhD students without GPU clusters. Open base models = real SFT/RL research, not just prompting chat models. Respect to
Alibaba shipped four Qwen 3.5 small models with a trick borrowed from their 397B model: Gated DeltaNet hybrid attention. Three layers of linear attention for every one layer of full attention. The linear layers handle routine computation with constant memory use. The full
The Qwen 3.5 small model hype is getting ahead of itself. Yes, the 9B beats GPT-5 Nano by 13 points on MMMU-Pro (70.1 vs 57.2) and 30+ points on document understanding. Yes, it outperforms Qwen's own previous-gen 30B on most benchmarks at a third the size. The bar charts look