On September 23, the Qwen large model launched the Qwen-Audio-3.1 series of audio large models, forming a complete audio capability stack covering understanding, generation, interaction, and creation. The series includes speech recognition Qwen-Audio-3.1-ASR, audio understanding ASR-Next, text-to-speech TTS, audio creation TTS-Next, and real-time interaction Realtime. The API has been listed on the Qwen AI platform, and the ASR-Next API will be launched soon. At the same time, all products are being discounted, with TTS dropping by about 70%, Realtime by about 85%, and ASR by up to 95%.

In terms of capabilities, ASR enhances multilingual and dialect recognition as well as context understanding. It now supports native transcription polishing, automatically removing filler words and repeated expressions, and supports end-to-end role-based transcription, jointly outputting speaker tags, timestamps, and text; one model supports 30 languages and 16 Chinese dialects, with a first-word response time of about 160 milliseconds for streaming recognition.
ASR-Next is based on the next-generation architecture, expanding audio understanding from text in speech to emotions, ambient sounds, and mechanical sounds, supporting voice description, event localization, and audio question-answering reasoning. TTS supports natural cross-language voice color migration and controls emotion and speaking speed through instructions. TTS-Next is aimed at audio creation, unifying the generation of human voice, sound effects, and ambient sounds, supporting multi-role voice replication, fine-grained timestamps, and 48kHz output, serving scenarios such as podcasts, audiobooks, and film and game applications. Realtime realizes full-duplex interaction based on a multi-teacher distillation architecture, allowing interruptions at any time, supporting real-time language switching, emotion understanding, and tool calls during dialogue.
In testing, ASR achieved an average CER of 4.55% on open-source dialect test sets, an average CER of 10.38% on self-built Chinese dialect sets, and an average semantic sentence accuracy of 82.10% for converting dialects to Mandarin; ASR-Flash-Next and ASR-Flash ranked first in five and three out of eight metrics across four test sets, comprehensively outperforming the previous Fun-ASR cascading system.
Currently, Realtime is integrated with agents like Qoder and Qwen Office, as well as smart hardware like the Qwen AI glasses, continuing the trend of voice becoming a natural interface for human-computer interaction.
Qwen-Audio-3.1-ASR:
https://www.qianwenai.com/models/qwen-audio-3.1-asr-flash
Qwen-Audio-3.1-TTS:
https://www.qianwenai.com/models/qwen-audio-3.1-tts-flash
Qwen-Audio-3.1-TTS-Next:
https://www.qianwenai.com/models/qwen-audio-3.1-tts-next
Qwen-Audio-3.1-Realtime:
https://www.qianwenai.com/models/qwen-audio-3.1-realtime-plus
(The Qwen-Audio-3.1-ASR-Next API will be launched soon)
Join Now