Sunday, Sep 27, 2026
1 Xiaomi Open-Sources 7,000-Plus MiMo RL Environments and Training Code on Hugging Face 🤖 AI Sep 24, 11:00 AM EDT 89/51
Xiaomi released roughly 7,000 reinforcement learning task environments used to train its MiMo models on Hugging Face, spanning code, cybersecurity, general tasks, music and web development, along with the end-to-end RL training framework. The environments ship in a uniform Harbor format that lets users pick a task, a model, a harness and a sandbox provider. Xiaomi also open-sourced a Qwen model distilled from MiMo RL trajectories as a starting point for further RL work.
Hugging Face CEO Clement Delangue called the release a landmark for frontier open-source RL environment data, noting that comparable tasks usually cost hundreds to thousands of dollars each, which makes the repo "literally worth millions." One reviewer pointed out the repo also contains a smaller selection of 989 environments used to train the 9B distilled model rather than the flagship MiMo.
The environments are the ones behind MiMo-V2.6-Pro, the top open-weight model on the Artificial Analysis Intelligence Index at 46, tying Grok 4.7 and beating GLM-5.3, at $0.435 per million input tokens. The RL run lifted Pro's DeepSWE score from 58.41 to 72.57, near the benchmark's high of 74, and cost $2.6M for Pro and $0.9M for Flash over roughly five days on a 10,000 GPU cluster. Xiaomi has since diagnosed and fixed a tool-call repetition issue in the MiMo-V2.6 series.