Qwen3.8-Omni-Flash, the new generation of native multimodal model, was launched on the 18th, supporting text, image, audio, and video input with a 1M-length context. It is now available for experience on the Qwen AI platform. Building upon the previous general Agentic capabilities in programming, text knowledge work, and GUI operations, this model focuses on expanding Agentic applications centered around audio and video, covering workflows such as video editing, MV creation, film production and commentary, audio/video to text summary, and audio/video conversation, which require comprehensive processing of multimodal content.

In a total of 30 evaluations, its average score increased by over 26% compared to the previous generation Qwen3.5-Omni-Plus. In areas such as Audio-Visual Agent, Coding, and long-term tasks, WildClawBench-MM improved by 36.5 points, AgenticVBench by 22.3 points, and UniClawBench achieved 69.6 points; on basic capabilities, LongAudioSpan improved by 8.3 points, OmniVideoBench by 9.6 points, and CSR/ISR of OmniCap-IF increased by 8.5/14.1 points respectively. The DER/cpWER of AliMeeting dropped from 88.11/89.61 to 3.35/17.18.

The official claims that its audio and video capabilities are close to Gemini3.8Flash, with overall audio capabilities exceeding it; the price per hour for audio input via API has decreased by more than 98%, and the price for audio and video input has dropped by more than 93%.
To support long-term workflows and real-time interaction, Qwen has expanded the Qwen-MM-Plugins and open-sourced Qwen-Live Harness. Under the Agentic Understanding mode, the accuracy of OmniVideoBench rose from 63.4 to 67.8, and the token consumption decreased from 145,736 to 79,117, a reduction of about 45.7%.

The team also brought the model into the R&D process itself: completing evaluation set selection, data construction, and four iterations within 12 hours, building a total of 3,413 training data entries. This reduced the character error rate of Sichuan dialect recognition for Qwen2.5-Omni-3B from 25.79% to 15.30%, a relative decrease of about 40.7%. The simultaneously released Realtime version is the first full-modal large model that supports "sound localization."
Join Now