On September 17, Xiaomi MiMo team lead Luo Fuli—formerly a DeepSeek researcher—publicly shared the latest reinforcement-learning training progress of the MiMo-V2.6 model with a live real-time page, marking her return to active technical competition after a six-month quiet period.
MiMo-V2.6 is currently in the middle of its RL phase; the team is scaling up compute, environment test frameworks and evaluation compute. By the afternoon of the 17th, cumulative training cost of the Pro and Flash versions had exceeded $1.35 million, at roughly $31,000 per hour—over 200,000 yuan an hour. The team plans to open-source related details in the coming weeks.
The live-stream is seen as a signal that Xiaomi's exploration of physical-world and general intelligence has entered a breakthrough phase, and that transparency around large-model training is becoming a new form of technical branding.
Source: AIBase / Xiaomi





