In preparation for the next generation of ultra-large models with rumored parameter scales exceeding 500 billion, ByteDance's Seed division has recently completed a deep restructuring of its organizational framework. This adjustment broke away from the traditional model of dividing teams by modality and research direction, re-integrating previously dispersed teams to comprehensively optimize R&D efficiency.

During this restructuring, the Seed Foundation Model officially established four new first-level departments. The first is the Pretrain Data team, led by Li Chenggang. This department was formed by merging pre-training teams that were previously spread across text, programming, visual understanding, and speech-related areas, and it will now be responsible for multi-modal data and large-scale model pre-training for the new Omni model.

The second is the Horizon RL team, led by Tang Shengyu. This team integrates post-training capabilities that were previously scattered across post-training, reasoning, and visual understanding areas, with the core goal of enhancing the basic intelligence ceiling of models through reinforcement learning.

For terminal and business-oriented post-training applications, ByteDance implemented more refined task breakdowns. The Product Posttrain-Work department, led by Qin Yujia, is responsible for B-end applications and the integration and release of Agentic models, focusing on optimizing the Agentic capabilities of models in office scenarios and supporting the task modes of Doubao and Dola. Meanwhile, the former Application team, which was previously led by Zhu Wenjia, has been renamed Product Posttrain-Chat, dedicated to the integration and release of C-end dialogue models, forming an efficient division of labor with the Work team. All four new departments report to Wu Yonghui, and Seed has also appointed separate leaders for frontier exploration areas such as AI safety.

Looking back, the Seed team at ByteDance used to be divided independently based on specific modalities and research directions. This mechanism could quickly enhance various capabilities during the early stages of model development, but as technical boundaries became blurred, cross-team communication costs increased, and issues like "reinventing the wheel" emerged. The underlying logic behind this restructuring is to "merge similar work, reduce overlap, and consolidate efforts."

Previously, media reports indicated that ByteDance is currently in the early stages of discussing the training of a model with a parameter scale exceeding 500 billion. If this plan is fully implemented, its parameter count would surpass Alibaba's Qwen 3.8-Max and Moonshot's K3, becoming the largest known model in terms of parameter scale in China. Facing this new phase of striving toward ultra-large models, this comprehensive organizational restructuring and division of labor undoubtedly provides a solid organizational guarantee for ByteDance in the intense global AI arms race.