On July 16, Xiaomi officially released the Xiaomi-Robotics-1, a embodied foundation model designed for real mobile operation tasks. The model was pre-trained on 100,000 hours of real-world data and completed training by combining cross-body data, marking a systematic step forward for Xiaomi in advancing embodied intelligence models along the "Scaling Law" (scale law) path.

QQ20260716-143733.jpg

Traditional robot strategy models are often limited by hardware dependencies and scarce data scales. To break through this bottleneck, the Xiaomi team introduced 100,000 hours of real-world trajectories collected through the UMI (Universal Manipulation Interface) device during the pre-training phase, covering multiple scenarios such as home, commercial, and industrial environments, and combined it with an efficient visual language model to complete full-scale automatic annotation within two weeks.

QQ20260716-143752.jpg

In the post-training phase, the team used approximately 10,000 hours of cross-body data for body and instruction alignment, enabling Xiaomi-Robotics-1 to have "out-of-the-box" multi-type mobile operation capabilities. Experiments show that as the amount of training data and model size (offering three versions: 2B, 5B, and 10B) increases, the model demonstrates clear scaling growth trends in action prediction accuracy and success rate of unseen scenario tasks, and has set new SOTA (state-of-the-art) records on multiple public simulation benchmarks such as RoboCasa365 and RoboDojo.

This release of Xiaomi-Robotics-1 not only demonstrates Xiaomi's R&D strength in the field of physical AI but also successfully verifies a scalable embodied intelligence training path: "large-scale pre-training - cross-body post-training - fine-tuning with a small amount of data," providing a highly referenceable paradigm for robots to transition from lab demonstrations to complex and realistic physical worlds.