Alibaba's new Qwen3.5-Omni multimodal model, which processes text, audio, images, and video, is proprietary, marking a shift away from its open-source strategy
Alibaba has built the Qwen line through repeated open releases, from its edge-deployable Qwen2.5-Omni model to open-source Qwen3-Omni models that handled the same broad mix of media. The proprietary status of this release is therefore a meaningful break in how Alibaba distributes its flagship multimodal capability.
The change also lands as Alibaba seeks to monetize AI and has raised prices for AI chips and cloud storage amid stronger demand. Holding this model closed gives the company more control over how a higher-end model is accessed and commercialized.
First-order effects
Developers and enterprises that expected open model weights from the Qwen Omni line must use Alibaba-controlled access rather than independently deploy or modify this release.
Alibaba gains direct control over distribution, usage terms and potential monetization of its latest multimodal model.
Second-order effects
The move makes Alibaba’s cloud and model-access offering more central to customers that want its newest multimodal capabilities, reinforcing its broader AI monetization effort.
Open-model users may weigh older Qwen releases against proprietary alternatives on flexibility and access, while competing model providers face a clearer split between open distribution and controlled services.
Third-order effects
If Alibaba continues to reserve leading releases, Qwen could evolve into a tiered portfolio: open models for adoption and ecosystem reach, proprietary models for premium capabilities and revenue capture.
The shift is one data point in a broader contest over whether multimodal AI value accrues primarily to open-weight ecosystems or to providers that control model access and distribution.
The trend: Frontier multimodal AI is increasingly being positioned as a controlled commercial service even by companies that used open releases to build developer adoption.
🚀 Qwen3.5-Omni is here! Scaling up to a native omni-modal AGI. Meet the next generation of Qwen, designed for native text, image, audio, and video understanding, with major advances in both intelligence and real-time interaction. A standout feature: ‘Audio-Visual Vibe Coding’. [i…
Alibaba's Qwen3.5-Omni just dropped with script-level captioning, audio-visual vibe coding, and real-time web search built in. However, there is a catch: Omni here doesn't mean *creating* image or voice, but rather interpreting it. So, a caveat. Open access via Hugging. [image]
Qwen3.5-Omni might be the strongest multimodal frontier model right now. What impressed me most: audio-visual vibe coding. Point your camera at something, describe what you want, and it turns that into working code. Really hope this gets open-sourced soon.
🚀 Introducing Qwen3.5-Omni, the latest fully omnimodal LLM in the family. With exceptional full-modality perception and generation capabilities, it's built to drive the next generation of AI applications. #AlibabaAI #Qwen
Qwen @Alibaba_Qwen just released Qwen3.5-Omni 🔥 Weights are not released ( yet?), but you can try the demos: ✨ Online demo https://huggingface.co/... ✨ Offline demo https://huggingface.co/...
1/10 🚀 Qwen3.5-Omni is here! Scaling up to a native omni-modal AGI. Meet the next generation of Qwen, designed for native text, image, audio, and video understanding, with major advances in both intelligence and real-time interaction. A standout feature: Audio-Visual Vibe [image]