/
Navigation
Chronicles
Browse all articles
Explore
Semantic exploration
Research
Entity momentum
Nexus
Correlations & relationships
Story Arc
Topic evolution
Drift Map
Semantic trajectory animation
Posts
Analysis & commentary
Pulse API
Tech news intelligence API
Browse
Entities
Companies, people, products, technologies
Domains
Browse by publication source
Handles
Browse by social media handle
Detection
Concept Search
Semantic similarity search
High Impact Stories
Top coverage by position
Sentiment Analysis
Positive/negative coverage
Anomaly Detection
Unusual coverage patterns
Analysis
Rivalry Report
Compare two entities head-to-head
Semantic Pivots
Narrative discontinuities
Crisis Response
Event recovery patterns
Connected
Search: /
Command: ⌘K
Embeddings: large
TEXXR

Chronicles

The story behind the story

days · browse · Enter similar · o open

Alibaba debuts its Qwen3 family of open-weight “hybrid” AI reasoning models, including Qwen3-235B-A22B, with 235B total parameters and 22B activated parameters

Chinese tech company Alibaba on Monday released Qwen3, a family of AI models the company claims matches …

TechCrunch Kyle Wiggers

Context & Ripple Effects

This release established the Qwen3 line’s open-weight, hybrid-reasoning foundation. It was followed by an instruction-tuned Qwen3-235B-A22B update focused on reasoning, accuracy, and multilingual capability, indicating that Alibaba treated the initial family as a continuing model platform rather than a one-off launch.

The subsequent Qwen releases broadened that platform in both directions: smaller open-weight Qwen3.5 variants extended the line to lighter deployments, while later previews pushed toward much larger models. That makes the original architecture and release posture consequential beyond a single benchmark cycle.

First-order effects

  • Developers and enterprises gain an open-weight Qwen option centered on reasoning, with a 235B-total-parameter model that activates 22B parameters; this creates a concrete alternative for teams able to run and adapt their own models.
  • Alibaba gains a reusable Qwen3 base for follow-on releases and distribution, rather than relying solely on a closed hosted-model offering.

Second-order effects

  • Model buyers can weigh total capability against active-parameter requirements when assessing self-hosted deployments, increasing pressure on rival open-weight providers to show practical efficiency as well as headline model scale.
  • The release creates a foundation for specialized derivatives, as illustrated by the later Qwen3 instruction-tuned release, shifting more differentiation toward tuning, tooling, and deployment support.

Third-order effects

  • If this pattern persists, open-weight frontier-adjacent models will make the base model less of a standalone moat and place greater value on complements such as infrastructure, customization, and distribution.
  • The Qwen roadmap—from small variants to very large previews—suggests a bifurcated model market: broadly deployable open-weight options alongside selectively opened high-end systems, with the balance determined by operating cost and ecosystem uptake.

The trend: This is one data point in the industrialization of AI models, where open-weight releases and efficient activation architectures expand buyer choice while shifting competition to the surrounding stack.

Discussion

  • @teknium1 @teknium1 on x
    Idk why Americans don't just embrace Apache and MIT but it's a big deal
  • @justinlin610 @justinlin610 on x
    Qwen3 is finally out! It really takes some time for our guys to figure out methods to solve some problems that are not fancy. How to scale RL with stable training, how to balance data from different domains, how to increase the support of more languages with performance
  • @alibaba_qwen @alibaba_qwen on x
    Introducing Qwen3! We release and open-weight Qwen3, our latest large language models, including 2 MoE models and 6 dense models, ranging from 0.6B to 235B. Our flagship model, Qwen3-235B-A22B, achieves competitive results in benchmark evaluations of coding, math, general [image]
  • @unsorsodicorda Andrea Panizza on x
    Amazing models. I'll be a contrarian and while everyone will (understandably!) raving about Qwen3-235B-A22B, I'm mostly shocked by Qwen3-32B (incredible scores for a 32B model, I think this literally punches the Pareto front through the roof!), Qwen3-30B-A3B (it looks much 1/
  • @dorialexander Alexander Doria on x
    First tests on qwen 0.6b does confirm tiny models could overcome their memory issues through recursive reasoning *but* we need to better design reasoning processes to reopen generation paths. Right now it's running circle on the first assumptions. [image]
  • @iscienceluvr Tanishq Mathew Abraham, Ph.D. on x
    The Qwen team really cooked with this release Just incredible work all around: - 235B MoE that is comparable to o1, o3-mini, Gemini 2.5 Pro, etc. - trained on 36T tokens, covering 119 languages! Data extracted from PDFs, synthetic data, etc. - Thinking and non-thinking modes - [i…
  • @alibaba_qwen @alibaba_qwen on x
    Qwen3 models are supporting 119 languages and dialects. This extensive multilingual capability opens up new possibilities for international applications, enabling users worldwide to benefit from the power of these models. [image]
  • @mascobot @mascobot on x
    Wow, Qwen3 235B MoE with 22B active params, beats R1, Grok, O1 and O3 mini. - 2 MoE models and 6 dense models, ranging from 0.6B to 235B. - Apache 2.0 [image]
  • @natolambert Nathan Lambert on x
    Have fun at Qwen3 ̶l̶l̶a̶m̶a̶ Con tomorrow!
  • @ns123abc Nik on x
    > Qwen3 drops > 235B total / 22B active > neck to neck with LLaMA-4-Maverick zucc in absolute shambles [image]
  • @teortaxestex @teortaxestex on x
    Qwen's distillation teacher MoE has fewer total parameters (235B) than Meta's Llama 4 Behemoth has active (288B). As a result its still much smaller distills dunk on Scout viciously. Should have trained that behemoth ass to a fit condition first, huh [image]
  • @xenovacom @xenovacom on x
    I seriously cannot believe this is a 0.6B LLM! 🤯 @Alibaba_Qwen just released Qwen3, a series of hybrid reasoning models that allow you to control how much “thinking” the model does for a given task. They can even run locally in your browser on WebGPU with 🤗 Transformers.js! [vide…
  • @alibaba_qwen @alibaba_qwen on x
    Qwen3 exhibits scalable and smooth performance improvements that are directly correlated with the computational reasoning budget allocated. This design enables users to configure task-specific budgets with greater ease, achieving a more optimal balance between cost efficiency and…
  • @kaggle @kaggle on x
    🤖 Now on #KaggleModels! The long-awaited @Alibaba_Qwen's Qwen 3 is here - featuring dense + MoE models, trained on 36T tokens across 119 languages. Big gains in reasoning, performance, and long-context (32k tokens) support! Learn more: https://www.kaggle.com/...
  • @theo @theo on x
    Qwen3 maintains the Qwen trend of massively overthinking tasks, generating thousands of thinking tokens and running out of context before answering. [video]
  • @alibaba_qwen @alibaba_qwen on x
    We have optimized the Qwen3 models for coding and agentic capabilities, and also we have strengthened the support of MCP as well. Below we provide examples to show how Qwen3 thinks and interacts with the environment. [video]
  • @huybery @huybery on x
    The long-awaited Qwen3 is finally here! Our team has put tremendous effort into Qwen3, hoping to bring something fresh to the open LLM community. We've made significant progress in pretraining, large-scale reinforcement learning, and integration of reasoning modes. We believe
  • @ollama @ollama on x
    ollama run qwen3 You can toggle non-thinking mode in Ollama by typing /no_think after the prompt. Check out all the model sizes 👇👇👇🧵
  • @nisten @nisten on x
    qwen means business apache license too even the smallest tinies 600mb weight is crazy lol, the 0.6B I tested got 38% on medmcqa medical lol [image]
  • @bartowski1182 Colin Kealty on x
    Qwen 3 is live!!! The legends at @Alibaba_Qwen have graced us with 8 new model sizes, including some nice small ones, a dense 32B, and an MoE 30B as well as a 235B MoE! Everything but the 235B is ready to go on @lmstudio: https://huggingface.co/... As well as on my own page!! 🤗
  • r/singularity r on reddit
    Qwen3: Think Deeper, Act Faster
  • r/LocalLLaMA r on reddit
    Qwen3: Think Deeper, Act Faster