Lead story
Models & availability
Latest
Lead story
Models & availability
Latest
Alibaba's Qwen announces Qwen3.5-Omni, a new fully omnimodal LLM supporting text, images, audio, and audio-visual content. It comes in three sizes: Plus, Flash, and Light, with 256k long-context input. It is natively pretrained on massive data and offers enhanced multilingual capabilities compared to Qwen3-Omni. The model is available via Offline and Realtime APIs, and features include audio-visual captioning, voice cloning, and ARIA technique for speech stability.
From the source
Qwen3.5-Omni is Qwen’s latest generation of fully omnimodal LLM, supporting the understanding of text, images, audio, and audio-visual content. Both the Thinker and Talker in Qwen3.5-Omni adopt the Hybrid-Attention MoE. Qwen3.5-Omni series includes Instruct versions in three sizes: Plus, Flash, and Light, with support for 256k long-context input.
qwen.ai