/
Navigation
Chronicles
Browse all articles
Explore
Semantic exploration
Research
Entity momentum
Nexus
Correlations & relationships
Story Arc
Topic evolution
Drift Map
Semantic trajectory animation
Posts
Analysis & commentary
Pulse API
Tech news intelligence API
Browse
Entities
Companies, people, products, technologies
Domains
Browse by publication source
Handles
Browse by social media handle
Detection
Concept Search
Semantic similarity search
High Impact Stories
Top coverage by position
Sentiment Analysis
Positive/negative coverage
Anomaly Detection
Unusual coverage patterns
Analysis
Rivalry Report
Compare two entities head-to-head
Semantic Pivots
Narrative discontinuities
Crisis Response
Event recovery patterns
Connected
Search: /
Command: ⌘K
Embeddings: large
TEXXR

Chronicles

The story behind the story

days · browse · Enter similar · o open

Alibaba releases its Qwen3.5-Omni omnimodal LLM with support for 10+ hours of audio input, saying the Plus variant surpasses Gemini 3.1 Pro on audio benchmarks

Qwen3.5-Omni is Qwen's latest generation of fully omnimodal LLM, supporting the understanding of text, images, audio, and audio-visual content.

Qwen

Context & Ripple Effects

Qwen3.5-Omni extends Alibaba’s omnimodal line after the open-source Qwen3-Omni family established text, image, audio and video processing as a single product category. The new model raises the practical ceiling for long-form audio analysis while keeping those modalities together.

The release also coincides with a reported move to make Qwen3.5-Omni proprietary, a notable departure from the earlier open-source positioning. That makes the model both a capability upgrade and a test of whether Alibaba will reserve its newest multimodal advances for controlled access.

First-order effects

  • Alibaba gains a new flagship multimodal model for workloads involving lengthy recordings and combined audio-visual inputs.
  • The claimed audio-benchmark advantage puts Gemini 3.1 Pro directly in the comparison set for buyers evaluating high-end audio understanding.

Second-order effects

  • Competing multimodal vendors face added pressure to demonstrate long-context audio quality, not just text, image or short-clip performance.
  • Developers choosing between Qwen and Gemini gain a more explicit basis for routing audio-heavy workloads, though benchmark claims still require task-specific validation.

Third-order effects

  • If leading models continue to compete on sustained audio and audio-visual understanding, multimodal capability will become a core platform requirement for ambient and voice-driven software rather than a standalone feature.
  • Alibaba’s apparent shift from open releases toward proprietary flagship models could sharpen the divide between models used for ecosystem adoption and models reserved for premium differentiation.

The trend: Frontier AI competition is expanding from general multimodality toward models that can process long-duration, continuous real-world audio and video inputs.

Discussion

  • @ali_tongyilab @ali_tongyilab on x
    1/10 🚀 Qwen3.5-Omni is here! Scaling up to a native omni-modal AGI. Meet the next generation of Qwen, designed for native text, image, audio, and video understanding, with major advances in both intelligence and real-time interaction. A standout feature: Audio-Visual Vibe [image]
  • @adinayakup Adina Yakup on x
    Qwen @Alibaba_Qwen just released Qwen3.5-Omni 🔥 Weights are not released ( yet?), but you can try the demos: ✨ Online demo https://huggingface.co/... ✨ Offline demo https://huggingface.co/...
  • @alibaba_qwen @alibaba_qwen on x
    Demo1:Audio-Visual Captioning [video]
  • @kimmonismus @kimmonismus on x
    Alibaba's Qwen3.5-Omni just dropped with script-level captioning, audio-visual vibe coding, and real-time web search built in. However, there is a catch: Omni here doesn't mean *creating* image or voice, but rather interpreting it. So, a caveat. Open access via Hugging. [image]
  • @alibaba_qwen @alibaba_qwen on x
    Demo2:Audio-Visual Vibe Coding [video]
  • @alibaba_qwen @alibaba_qwen on x
    🚀 Qwen3.5-Omni is here! Scaling up to a native omni-modal AGI. Meet the next generation of Qwen, designed for native text, image, audio, and video understanding, with major advances in both intelligence and real-time interaction. A standout feature: ‘Audio-Visual Vibe Coding’. [i…
  • @alibabagroup @alibabagroup on x
    🚀 Introducing Qwen3.5-Omni, the latest fully omnimodal LLM in the family. With exceptional full-modality perception and generation capabilities, it's built to drive the next generation of AI applications. #AlibabaAI #Qwen
  • @bowang87 Bo Wang on x
    Qwen3.5-Omni might be the strongest multimodal frontier model right now. What impressed me most: audio-visual vibe coding. Point your camera at something, describe what you want, and it turns that into working code. Really hope this gets open-sourced soon.
  • r/singularity r on reddit
    Qwen3.5 Omni - Qwen's latest generation of fully omnimodal LLM