Shengshu Technology officially released the new Vidu S2 video large model series on September 15th. The newly released Vidu S2 includes two core models: Vidu S2-Avatar, which is aimed at generating continuous interactive digital characters, and Vidu S2-Editing, which is designed for real-time editing of input video streams. In addition, Shengshu Technology has further explored real-time spatial video generation and editing capabilities for VR headsets. The entire full-chain development was led by Zhang Jintao, a doctoral student of Professor Zhu Jun and the head of streaming video generation and inference at Shengshu Technology.

Regarding model features, Vidu S2-Avatar has achieved a leap from continuous voice control to higher interaction freedom. After users upload a character image, they can interact with the model through text or voice. The model not only supports expressions and postures during speaking but also enhances the ability to follow instructions for large movements such as dancing. More importantly, the model allows dynamic addition of items, clothing, or background reference images at any time during the video stream, enabling the character to pick up a cup, change into specified clothing, or enter a new scene. At the same time, the model introduces a visual feedback mechanism to ensure the reliability of continuous actions and states, such as "picking up a cup and then smiling." Additionally, S2-Avatar has increased the real-time output resolution from 540P in the previous generation to 720P.

image.png

Vidu S2-Editing focuses on receiving continuous video streams and performing real-time editing based on text instructions and optional reference images, allowing the visuals to naturally change with actions such as raising a hand, turning around, or camera movement. Currently, this model mainly supports four types of tasks: first, style transfer, which changes the brushwork and color of the image while maintaining the person's contour, posture, and scene layout; second, virtual try-on, transferring the reference clothing onto the video character, accurately handling the clothing boundary, material texture, and occlusion relationship with the body; third, character replacement, transferring the face, hairstyle, clothing, and appearance of the reference character to the source video while retaining the original posture and movement; fourth, background replacement, changing the environment according to the reference image while maintaining stable spatial relationships when the person moves continuously.

Regarding spatial video exploration, Vidu S2 integrates the real-time output of Avatar into the spatial video conversion process, generating synchronized left and right eye views. Editing also supports first editing a regular monocular video and then converting it into a spatial video, or directly editing an existing spatial video. Both modes support the above four types of editing tasks, and the results can be streamed to VR headsets.

image.png

In terms of underlying technology, Vidu S2 has made several key technological innovations. First, it proposed the Self-Replay Forcing (SRF) training method. The model first generates longer self-generated trajectories and then samples continuous segments from them, re-noises them, and performs causal replay training with propagatable gradients, effectively avoiding identity drift or action breaks caused by small errors accumulating. Second, in data processing, it balances the richness of motion in dance videos and the stability of the image, using event sequence descriptions for video annotations and preference reward optimization. At the same time, S2-Avatar introduced a VLM agent as visual feedback to check the completion of actions and adjust subsequent prompts. In terms of performance and efficiency optimization, it uses a lightweight single-step Refiner to restore high-resolution details, combines TurboDiffusion and TurboServe technologies to reduce computational costs, and achieves shared timeline scheduling across modules.

In public and professional evaluations, Vidu S2 demonstrated excellent comprehensive performance. In the StreamAV-Bench evaluation, S2-Avatar achieved the best results in nine indicators, including visual and audio quality, audio-video synchronization, and consistency between the main subject and background. In video editing, S2-Editing scored 4.26 in the joint evaluation of OpenVE and RefVIE, and achieved the best results in all indicators on Sparkle-Bench. It also significantly surpassed competing models in the ViViD virtual try-on test and various professional human preference comparisons.

Regarding future plans, Shengshu Technology stated that spatial video is still in the stage of fixed perspective exploration. How to maintain clear image quality in headsets and further reduce latency, as well as exploring panoramic spatial videos, will be the core challenges ahead.