Qwen-Audio-3.1 series of speech large models was officially released today. This upgrade not only comprehensively enhances three core models: speech recognition, speech synthesis, and real-time speech interaction, but also introduces two new audio creation models: Qwen-Audio-3.1-TTS-Next and Qwen-Audio-3.1-ASR-Next. Five new models are launched at the same time, forming a complete audio capability stack covering "understanding - generation - interaction - creation." To further reduce the usage barrier, the prices of all Qwen-Audio voice models have been reduced, with TTS down by about 70%, Realtime down by about 85%, and ASR down by as much as 95%.

Five Models Launched Simultaneously, ASR Price Drops 95%
Qwen-Audio-3.1-ASR is a next-generation speech recognition model that adds native transcription polishing capabilities, automatically removing filler words and repeated expressions and reorganizing semantics. Its role-based transcription can fully present "who said what and when," providing traceable structured records for meetings and interviews. The model supports 30 languages and 16 Chinese dialects, industry terms, and graded hot word recognition, and provides low-latency streaming recognition, with a first character response time of about 160 milliseconds. In performance, its average character error rate (CER) on 11 subsets of KeSpeech and WSYue is 4.55%, an average CER of 10.38% on internal Chinese dialect test sets, and an average semantic sentence accuracy of 82.10% on AST.
The ASR-Next based on the next-generation architecture expands the processing object from "text in speech" to "complete audio information," understanding emotions, ambient sounds, and mechanical sounds, and completing sound descriptions, event positioning, and audio question-answering reasoning. In the role-based ASR evaluation, the Flash-Next and Flash versions achieved five and three best results respectively, surpassing the previous Fun-ASR level system comprehensively.
In speech synthesis, Qwen-Audio-3.1-TTS supports multilingual and dialect synthesis, allowing the same voice to naturally migrate across languages, and users can control emotion, speed, and expression through instructions. The new TTS-Next goes even further, generating human voices, sound effects, and ambient sounds by unifying text, timestamps, and reference audio, creating complete audio content in one go, and supporting 48kHz output and fine-grained timestamp control, meeting the needs of podcasts, audiobooks, films, and games.

From Understanding to Creation, Speech Becomes a Natural AI Entry Point
The real-time interaction model Qwen-Audio-3.1-Realtime uses a full-duplex architecture, allowing simultaneous speaking and listening, supporting interruptions at any time, and upgrading from text recognition to understanding emotions and communication intent—when it detects a user's low mood, it slows down the speech rate and adjusts the wording. It also supports real-time multi-language switching and completes tasks by calling APIs, knowledge bases, and business system tools during conversations; based on a multi-teacher distillation architecture, it maintains low latency while enhancing factual reliability and safety boundaries.
Join Now