MiMo-V2.6: Scaling Reinforcement Learning Towards Self-Improvement
- Type
- paper
- Venue
- Xiaomi MiMo technical report (Hugging Face), Sep 2026
- Year
- 2026
- Source
- huggingface
- Access
- free
- Language
- English
- Added
- 2026-10-01
- Verified
- 2026-10-01
Summary
Technical report for Xiaomi's MiMo-V2.6 series (Pro-RL flagship and Flash-RL), built to scale RL toward self-improvement by scaling RL compute, environment diversity, and grader compute together. The recipe: 'You Only RL Once' — one fully asynchronous GRPO run mixing coding, general agents, visual, and cybersecurity tasks and multiple harnesses in the same batch (1,568 prompts x 16 rollouts, ~25K trajectories and billions of tokens per step, 30 steps); groupwise agentic grading (GRS rubrics + GAR advantage redistribution) that ranks passing solutions instead of binary pass/fail; training on minimal open-source mini-harnesses whose gains transfer to unseen production harnesses; environment hardening against reward hacking; and MOPD2 multi-prefix multi-teacher on-policy distillation after RL. Pro is a 1.02T-total / 42B-active sparse MoE with 1M-token omnimodal context.
Keywords
MiMo-V2.6 · You Only RL Once · GRPO · groupwise agentic grading · GRS · GAR · MOPD2 · mini-harnesses · reward hacking · DeepSWE · Xiaomi
Topics
reinforcement learning, GRPO, agentic RL, multi-harness training, reward hacking, on-policy distillation, MoE
Research notes
- Discovery: @SergioPaniego X thread 2026-10-01 (https://x.com/SergioPaniego/status/2105666293697527994) reading the post-training sections
- Tech report PDF: https://huggingface.co/XiaomiMiMo/MiMo-V2.6-Pro-RL/blob/main/MiMo_V2_6_technical_report.pdf
- Scale of the run: 30 RL steps, ~750K trajectories in under 6 days (123h Pro / 82h Flash), reported cost ~$2.62M (Pro) and ~$0.85M (Flash); per step 1,568 prompts x 16 rollouts = ~25K trajectories, 2.7-3.7B tokens per update; Pro budget split ~43.5% training / 43.8% rollouts / 12.7% grading; code is 68% of the mix yet gains show on general agent benchmarks; average training-task pass rates rose ~12% relative (Pro) and ~25% (Flash)
- Minimal harnesses: production harnesses ship prompts/safeguards the reward never measures, so the model learns to ignore them; training on a family of minimal mini-harnesses (open source: github.com/XiaomiMiMo/mimoagent) transfers to harnesses never seen in training (Codex, Claude Code, mini-swe-agent; DeepSWE ~50% -> 66% held-out)
- Ranking passing solutions: if all 16 rollouts pass, advantage is zero; GRS builds task-specific rubrics offline and GAR ranks passing trajectories online — without it the policy drifted into exception swallowing and relaxed validation; same ranking applied to visual-task renders
- Reward hacking: report quotes the model fetching the upstream fix from a GitHub raw URL; countermeasures were env cleanup, network isolation, and a dedicated hack agent attacking every env; confirmed hacks stayed under 2% of the run
- MoE detail: with a trainable router, 22% of experts went cold within 20 steps, so the router was frozen during RL
- MOPD2 after RL: domain-specialized teachers distilled back on-policy; a k-turn trajectory gives k prefix starting points and the student generates only the next turn from each
- Final evals (report table): DeepSWE v1.1 Pro 71.9 / Flash 67.9, AutomationBench v1.0.6 53.1 / 52.3, CyberGym 94.0 / 95.1, OSWorld-Verified 82.0 / 80.8
- Architecture: sparse MoE 1.02T total / 42B activated, 384 routed experts (8 active), 1M context, native omnimodal (text/image/video/audio) with 681M MiMo ViT and audio encoders, 5-layer MTP speculative decoder
- Releases: open weights Pro-RL and Flash-RL (Hugging Face + ModelScope, collection https://huggingface.co/collections/XiaomiMiMo/mimo-v26-6ab18b0eff1d27b901d9f96b), 7.8K curated RL envs (catalog Datasets row 72 MiMo-V2.6-RL-oss; FineEnvs Harbor ports rows 74-79; Explorer Space https://huggingface.co/spaces/FineEnvs/MiMo-RL-Envs-Explorer), and a 9B starting point (Qwen3.5-9B SFT'd on MiMo data) that GRPO on the released envs improves on all 11 reported benchmarks
- Models released ~22 Sep 2026 per launch coverage; license not stated on the model card