Tencent Hunyuan has officially launched the new generation speech recognition model Hy ASR3.0 preview, which for the first time integrates the deep semantic understanding capabilities of the large language model Hy3 with high-precision speech recognition, driving speech recognition to evolve from "word-by-word transcription and single-point optimization" to "context understanding, scenario compatibility, and direct output."

According to official information, the new model performs outstandingly in multi-dimensional evaluations: on open-source evaluation sets, the word error rate (WER) for Mandarin Chinese, English, and Cantonese is as low as 3.34%, 2.62%, and 3.12%, respectively, with an overall rate around 3%; self-built evaluation sets cover general recognition, more than ten dialects, context semantic understanding, professional terminology recognition, and complex acoustic scenarios such as high noise and whispering, with WER maintaining industry-leading levels.

QQ20260804-163907.jpg

From a technical perspective, the capability upgrade of Hy ASR3.0 preview comes from full-chain optimization: the architecture adopts a MoE architecture that balances efficiency and performance, with the base upgraded to the Hy3 large model. It also includes a self-developed unsupervised speech encoder, trained on tens of millions of hours of speech data to enhance acoustic feature extraction. During the pre-training phase, the speech encoder and large language model are jointly trained, covering massive data with multiple dialects, accents, and acoustic environments, and enhancing features like dialect recognition and context awareness through multi-stage capability injection. In the post-training phase, a SFT dataset covering more than 20 dialect regions was built, combined with multi-stage reinforcement learning to address issues like misrecognition and missed recognition in complex scenarios specifically.