← Back to live feed

Sunday, Sep 27, 2026

1
Xiaomi Open-Sources 7,000-Plus MiMo RL Environments and Training Code on Hugging Face
topics 🤖 AI tags AIAI ModelsAI Open SourceAI ReleasesAI Research keywords XiaomiClement Delangue

Xiaomi released roughly 7,000 reinforcement learning task environments used to train its MiMo models on Hugging Face, spanning code, cybersecurity, general tasks, music and web development, along with the end-to-end RL training framework. The environments ship in a uniform Harbor format that lets users pick a task, a model, a harness and a sandbox provider. Xiaomi also open-sourced a Qwen model distilled from MiMo RL trajectories as a starting point for further RL work.

Hugging Face CEO Clement Delangue called the release a landmark for frontier open-source RL environment data, noting that comparable tasks usually cost hundreds to thousands of dollars each, which makes the repo "literally worth millions." One reviewer pointed out the repo also contains a smaller selection of 989 environments used to train the 9B distilled model rather than the flagship MiMo.

The environments are the ones behind MiMo-V2.6-Pro, the top open-weight model on the Artificial Analysis Intelligence Index at 46, tying Grok 4.7 and beating GLM-5.3, at $0.435 per million input tokens. The RL run lifted Pro's DeepSWE score from 58.41 to 72.57, near the benchmark's high of 74, and cost $2.6M for Pro and $0.9M for Flash over roughly five days on a 10,000 GPU cluster. Xiaomi has since diagnosed and fixed a tool-call repetition issue in the MiMo-V2.6 series.

Image via @iscienceluvr on X
Continues from Thursday, Sep 24
Xiaomi MiMo-V2.6 Takes No. 1 Open Weight Rank on Vals Index
79 tweets • 46 sources
See all 89 tweets →