Recently, the Tencent Hunyuan multimodal team has undergone a deep personnel restructuring and strategic adjustment. Lu Xudong, the former head of xAI multimodal understanding, has officially joined Tencent Hunyuan as the head of multimodal content generation algorithms. This change is accompanied by a series of intensive adjustments, including the departure of the former multimodal understanding head Hu Han to start his own business and the joining of Tian Yonglong, a former researcher at OpenAI. It also marks a profound reconfiguration in Tencent's multimodal large model development strategy.
Looking back at the history of Tencent Hunyuan's multimodal development, its technological landscape mainly revolved around several core products. Among them, the text-to-video model HunyuanVideo was launched in late 2024 as the largest open-source video generation model, followed by a series of expansions in image-to-video, customized generation, and digital human driving capabilities in 2025. In terms of image generation, from the early HunyuanImage to the 2.1 version supporting native 2K resolution and the 3.0 version with 80 billion parameters, it has continuously evolved toward an industrial level. Additionally, Hy3D, a single 3D model, has been widely applied in e-commerce modeling and product design. In the field of spatial intelligence, Tencent has also launched the world's first open-source, simulation-supporting immersive 3D world generation model, Hy World 1.0, and its subsequent versions, which deeply align with the direction of spatial intelligence research.
However, as Tencent merged the large language model department and the multimodal model department into the "Basic Model Department" and placed it under the full control of Yao Shunyu, the company's multimodal R&D approach is undergoing a profound transformation. In the current technical roadmap, there are different views in the industry on the relationship between multimodal and intelligent mainline. World Labs, led by Fei-Fei Li, believes that world models, as part of spatial intelligence, are an important mainline for AGI, while DeepSeek advocates using visual modalities as tools serving language models, weakening the independent 3D or world model approach.
From Yao Shunyu's recent strategic layout and the launch of the flagship product Hy3, it seems he leans more towards the latter technical path. Yao Shunyu has repeatedly emphasized in public that the competitive barrier in the second half of AI lies in context and reasoning abilities, not just the number of model parameters. By deeply integrating multimodal understanding capabilities into the basic model mainline, optimizing tool calling stability and long-context support, Tencent is committed to improving the understanding and execution efficiency of models in real-world office scenarios and GUI Agent products. The addition of Lu Xudong aims to enhance multimodal understanding capabilities, further addressing pain points such as inconsistency and lack of spatial stability in generation models. This restructuring reflects the alternation of internal technical strategies within Tencent and demonstrates its firm determination to align with the mainline and focus on productivity in the second half of the AGI competition.
Join Now