One-stop AI digital human desktop application AIGCPanel released version 2.5.0, bringing cloud-based voice capabilities that previously left users without a graphics card helpless directly to the desktop. The core of this update is straightforward: simply enter the API Key of Volcano Engine's Dobao Voice or Alibaba Cloud's DashScope in the "Other Models" section on the AI model page, and the cloud-based voice model can be immediately activated. People without any graphics card on their local machine can now complete high-quality voice synthesis through the cloud solution.
This step of entering the key has been made especially convenient. Before saving, the software will first synthesize a sentence "Hello, this is a voice preview." If the interface is truly connected, it allows saving, avoiding the need for users to later search error logs in tasks. Behind this is a voice provider registration table maintained by the main process: each vendor just needs to adapt its HTTP interface into a unified soundTts call to join. Dobao Voice uses X-Api-Key authentication, returning a concatenated stream of multiple JSON segments, and the software must parse and merge these fragments into a complete audio file; Alibaba Cloud uses DashScope's qwen-tts, returning an audio link or base64 data. In the future, adding a new vendor only requires writing one provider file.

The real change is that "voice" has been centralized into a manageable system from scattered configurations. On the digital human side, it is called "voice tone," while on the live streaming side, it is referred to as "live streaming tone." Each side has its own library, both supporting the addition of voice synthesis or voice cloning tones, with a direct preview before saving. These two sets of tones are stored separately and do not interfere with each other. For voice cloning, the reference audio needs to be re-recorded or re-uploaded, combined with corresponding text, no longer reusing the digital human's voice library—this kind of argument about "which user influenced the cloned voice" will no longer happen. On the right side of the add dialog box, there is also a helpful usage guide: recording requirements for cloning and operation guidance for synthesis.
At the bottom, there is a VoiceService factory with two storage identifiers SoundVoice and LiveVoice. Each voice record stores type, model identifier, parameters, and reference audio path. The real-time synthesis for preview and live streaming uses the same synthesize() call chain, with identical results. Even more cleverly, changing the voice during a live stream does not require modifying the live streaming service: the client only exposes one provider name with the Voice: prefix, and the rest of the routing is handled internally. In the future, adding a new voice only requires one line of code on the live streaming end. The live streaming configuration prioritizes the selected voice, and if none is selected, it falls back to the old voice driver configuration. Old users' previous configurations will not become invalid.
The live streaming control panel has also gained eyes. At the top, the new "Model Control" area allows starting or stopping the live streaming model and viewing the running status. Next to it, a "Model Log" button opens up the real-time log. If the live streaming model hasn't been imported yet, it will actively prompt and jump directly to the AI model interface to import it. Previously, confirming whether the "live streaming model was running" required switching between several pages, but now everything can be viewed in one area. It is implemented without building from scratch: filtering the model list to find records with live streaming capabilities, using existing components for status badges and log viewing.
Interface details have been thoroughly polished. Dialog boxes naturally expand in height, no longer overflowing the screen in small windows; the selected item in the left navigation changes to a light blue brand shade, and the dark mode is also adjusted to be brighter; disabled buttons are changed to light gray background with medium gray text, no longer making them hard to read. Several problematic bugs have also been fixed: general model tasks were once mistakenly marked as successful when the process exited abnormally, now they are verified to ensure actual output results; some error messages still displayed in Chinese on English interfaces have been fixed; errors caused by executing wmic on the new Windows version have also been eliminated.
Join Now