NetEase Youdao Launches Open Source Speech Synthesis Engine 'Yimosheng' Supporting Over 2000 Voice Tones


Qwen launches the Qwen-Audio-3.1 series of large speech models, with comprehensive upgrades in speech recognition, synthesis, and real-time interaction, and introduces the audio creation model TTS-Next and the understanding model ASR-Next. Five models cover the complete capability stack of 'understanding - generation - interaction - creation'. At the same time, all models are discounted, with TTS dropping by about 70% to lower the usage threshold.
The Gates Foundation will invest $1 billion over the next two years to promote the development of artificial intelligence, making cutting-edge technology benefit the poorest regions. Gates also shared his views on the global large model landscape, computing power carbon emissions, and Chinese humanoid robots, pointing out that according to different measures, there are currently four leading large model companies in both the US and China.
Alibaba has fully launched the Qwen-Audio-3.0 series of speech models on the Tongyi AI platform, including three types: speech recognition, speech synthesis, and real-time voice interaction. Officially, the series performed exceptionally well in the July 2024 speech ranking list by the authoritative evaluation platform Artificial Analysis, securing first place in all three categories: speech recognition, real-time interaction, and speech synthesis.
The Zhipu team has open-sourced four core video generation technologies, including GLM-4.6V visual understanding, AutoGLM device control, GLM-ASR speech recognition, and GLM-TTS speech synthesis models, showcasing their latest progress in the multimodal field and laying the foundation for the development of video generation technology.
In the context of rapid development in speech synthesis technology, Mianbi Intelligence and the Human-Computer Speech Interaction Laboratory at the Shenzhen International Graduate School, Tsinghua University (THUHCSI) recently jointly released a new speech generation model - VoxCPM. This model, with a parameter size of 0.5B, is dedicated to providing users with high-quality and natural speech synthesis experiences. The release of VoxCPM marks another milestone in the field of high-fidelity speech generation. The model has achieved industry-leading levels in key indicators such as naturalness, voice similarity, and prosodic expressiveness.