1. Returning to the Public Eye After Half a Year of Silence

The Xiaomi, which has been widely noticed in the field of artificial intelligence, has new developments. Luo Fuli, the head of the Xiaomi MiMo large model team and former researcher at DeepSeek, who had been silent for half a year, officially announced and live-streamed the current training process of the large model on the X platform at 12:00 AM on September 17th. This move marks her return to the intense competition in large model technology, and also allows the outside world to get their first glimpse into the internal training status of Xiaomi's latest generation model.

2. RL Experiments and Three Expansion Directions of MiMo-V2.6

According to the real-time training page shared by Luo Fuli, Xiaomi's latest large model, MiMo-V2.6, is currently conducting large-scale agent (intelligent entity) reinforcement learning (RL) experiments. In this run, the team focused on advancing three dimensions of technical expansion:

Computing resource expansion: Each step consumes about 2 billion tokens, including 1568 prompts multiplied by 16 rollouts, using a fully asynchronous operation mode.

Environment and testing framework expansion: A multi-task agent-based RL was built, which synchronously integrates multiple testing frameworks in a single run.

Evaluation computing expansion: Agent-based intra-group credit assignment was introduced, equipped with actual test cases and a reward mechanism based on benchmarks.

Luo Fuli also revealed that the team will gradually open-source related technical details to the public in the coming weeks.

3. High Training Costs and Performance of Two Versions

From the publicly available real-time training data, this large model training shows extremely high financial and computing power consumption. The training page currently mainly includes two parallel versions:

MiMo-V2.6-Pro: The cumulative training time reached 1 day and 19 hours, with a total cost of approximately $890,000, averaging more than $20,000 per hour of training.

MiMo-V2.6-Flash: The cumulative training time was 1 day and 14 hours, with a total cost of approximately $397,000, averaging over $10,000 per hour of training.

The total training costs of both versions have already exceeded $1.28 million, which translates to an overall hourly expenditure of over $30,000, fully highlighting the extreme demand for massive computing power in the reinforcement learning phase of top-tier large models.

4. Team Background and Previous Progress

In terms of business strategy, Xiaomi's senior management has repeatedly publicly disclosed phased achievements in the AI field. In March this year, Lei Jun, the founder of Xiaomi, mentioned on Weibo that the trillion-parameter large model MiMo-V2-Pro ranked eighth in the global comprehensive intelligence ranking of Artificial Analysis, and was fifth in brand ranking, surpassing xAI's Grok model and planned to maintain rapid iteration and enhancement afterward.

As a core leader, Luo Fuli is a highly acclaimed "95 post" AI young scientist. Public information shows that she graduated from the Computer Science major at Beijing Normal University, and then pursued a master's degree in computational linguistics at Peking University. After graduation, she joined Alibaba DAMO Academy's Machine Intelligence Lab as a researcher, responsible for the development of the multilingual pre-training model VECO, and promoted the open-source of the AliceMind project.