Tencent Hunyuan has officially launched the new generation speech recognition model Hy ASR3.0 Preview, marking a new stage in speech recognition technology from "word-by-word transcription" to "contextual understanding." The model is built on the latest generation large language model Hy3, integrating high-precision speech recognition and deep semantic understanding capabilities, with the core goal of achieving a qualitative leap—making AI not only clearly hear each word but truly understand what you are saying.

Supports ten dialect regions, with outstanding performance in long audio scenarios
In terms of coverage, Hy ASR3.0 Preview does not require standard Mandarin. It supports 10 major dialect regions, including Cantonese and Wu, and more than 20 secondary sub-regions. Open-source evaluation data shows that the model's word error rate (WER) for multiple languages is controlled at around 3%, with Mandarin at 3.34%, English at 2.62%, and Cantonese at 3.12%. The overall accuracy is already close to the industry ceiling.
Combined with Hy3's language understanding capabilities, the model can accurately capture user intent by considering the context, automatically eliminate ambiguities, and intelligently correct homophones. In scenarios such as meeting minutes and long audio interviews, this "understand first, then transcribe" approach shows significant advantages—traditional word-by-word transcription often struggles with homophonic ambiguities, while Hy ASR3.0 can automatically select the correct words based on semantic logic.
MoE architecture as the foundation, trained on millions of hours of data
On the technical level, the model uses an MoE architecture that balances efficiency and performance, with the base upgraded to Hy3. The Hunyuan team developed a self-researched unsupervised speech Encoder, which was trained on tens of millions of hours of unsupervised speech data to extract high-quality acoustic representations from complex audio—Encoder is responsible for "hearing well," while Hy3 is responsible for "understanding correctly."
In terms of data strategy, the team conducted joint training of the speech Encoder and large language model, introducing tens of millions of hours of multi-source speech data, and through multi-stage capability injection, enabling the model to have context awareness, adaptability to complex scenarios, and dialect recognition abilities. The post-training phase focuses on building a high-quality SFT solution covering general recognition, context understanding, and robustness in complex scenarios, covering context, specialized terms, different acoustic environments, and diverse voice inputs. It also introduces a multi-stage reinforcement learning approach to continuously optimize complex long-tail scenarios.
