From the source
Lead story
Top stories
Models & availability
Latest
Lead story
Top stories
Models & availability
Latest
From the source
Audio-visual input pricing drops 93%, model improves 25% across omnimodal benchmarks.
Alibaba released Qwen3.8-Omni-Flash, a native omnimodal model with a 1M-token context window, supporting text, image, audio, and video inputs.
The model improves average scores by over 25% compared to Qwen3.5-Omni-Plus across 29 evaluations, achieves audio-visual performance close to Gemini 3.8 Flash and overall audio performance exceeding Gemini 3.8 Flash, and reduces API prices for audio and audio-visual input by more than 98% and 93% respectively.
It also extends agentic capabilities to audio and video workflows such as video editing, music video creation, and real-time conversations.
From the source
Today, we are launching Qwen3.8-Omni-Flash , our next-generation native omnimodal model. ... Qwen3.8-Omni-Flash supports a 1M-token context window while maintaining text performance comparable to a text-only model of the same size and delivering significant improvements in omnimodal capabilities. Across 29 evaluations, its average score improves by more than 25% over Qwen3.5-Omni-Plus; the API price per hour of audio input decreases by more than 98%, and the price per hour of audio-visual input decreases by more than 93%.
qwen.ai