At 12:00 AM on September 17, Luo Fuli, head of the Xiaomi MiMo large model team and former DeepSeek researcher, posted on the X platform, publicly revealing for the first time the progress of the reinforcement learning (RL) training of the company's latest large model, MiMo-V2.6, and simultaneously launched a real-time training page, opening up a live broadcast of the entire RL phase of the large model. This move marks a new breakthrough period in Xiaomi's exploration of the physical world and general intelligence (AGI), with related technical details planned to be gradually open-sourced in the coming weeks.
In terms of technical architecture, MiMo-V2.6 is currently conducting large-scale multi-task agent (Agent) RL experiments. The team has comprehensively expanded in three dimensions: at the computing power level, it has achieved fully asynchronous parallelism for about 2 billion tokens per step (1568 prompts × 16 rollouts); at the environment level, it has built an agent-based RL system supporting mixed multiple frameworks in a single run; and at the evaluation level, it has adopted an internal credit allocation mechanism with test cases and benchmark rewards.
The live stream page displays the training progress, token consumption, and cost metrics in real time. The MiMo-V2.6-Pro version took approximately 1 day and 19 hours to train, costing $890,000, while the MiMo-V2.6-Flash version took 1 day and 14 hours, costing $397,000. The total training cost for both versions has exceeded $1.28 million.
As a "post-95" AI scientist previously brought in by Lei Jun from DeepSeek, Luo Fuli once led the development of multi-language pre-training at Alibaba DAMO Academy and DeepSeek-V2. She officially joined Xiaomi in November last year and took charge of the MiMo large model team. In March this year, the previous generation trillion-parameter model Mimo-V2-Pro ranked 8th on the Artificial Analysis Global Large Model Comprehensive Intelligence Ranking (5th in the brand ranking).
This MiMo-V2.6, through transparent live streaming and open-sourcing engineering details, demonstrates Xiaomi's deep accumulation in RL Scaling Law (Reinforcement Learning Scaling Law), and will also promote the evolution of cutting-edge large model training from a "black-box trial and error" approach to a "geek-level process open source" paradigm.
Join Now