Recently, Alibaba officially announced the full launch of the Qwen-Audio-3.0 series speech models on the Qwen AI platform. The series includes the large-scale speech recognition model Qwen-Audio-3.0-ASR, the large-scale speech synthesis model Qwen-Audio-3.0-TTS, and the real-time speech interaction dialogue model Qwen-Audio-3.0-Realtime.

According to the official introduction, the series models performed excellently in the voice rankings of the authoritative AI evaluation platform Artificial Analysis in July this year, winning first place in three core categories: speech recognition (ASR), real-time interaction (RealTime), and speech synthesis (TTS), achieving a "grand slam".

QQ20260817-155736.jpg

In terms of capabilities, Qwen-Audio-3.0-ASR supports identification in complex contexts and professional fields; Qwen-Audio-3.0-TTS supports multiple languages and dialects, and can accurately control emotions, tone, and rhythm; Qwen-Audio-3.0-Realtime realizes end-to-end speech understanding and dialogue, supporting interrupting and asking follow-up questions at any time, as well as tool calls, and has been applied to products such as Qwen App, Qwen Office, and Qoder.

Additionally, the Qwen AI platform has recently updated multiple main models, including Qwen3.8-Max, Qwen-image3.0pro, and Wan3.0, forming a comprehensive multi-modal model matrix covering text, code, speech, images, and video. The platform also launched a Token Plan subscription service with limited-time offers, starting at 39 yuan/month for the personal plan, and up to a 7.6 discount for the team plan, aiming to lower the access threshold for developers and create an AI productivity platform "born for Agents."