← Back to explorer

MiMo-V2.6: Scaling Reinforcement Learning Towards Self-Improvement

Type
paper
Venue
Xiaomi MiMo technical report (Hugging Face), Sep 2026
Year
2026
Source
huggingface
Access
free
Language
English
Added
2026-10-01
Verified
2026-10-01

Summary

Technical report for Xiaomi's MiMo-V2.6 series (Pro-RL flagship and Flash-RL), built to scale RL toward self-improvement by scaling RL compute, environment diversity, and grader compute together. The recipe: 'You Only RL Once' — one fully asynchronous GRPO run mixing coding, general agents, visual, and cybersecurity tasks and multiple harnesses in the same batch (1,568 prompts x 16 rollouts, ~25K trajectories and billions of tokens per step, 30 steps); groupwise agentic grading (GRS rubrics + GAR advantage redistribution) that ranks passing solutions instead of binary pass/fail; training on minimal open-source mini-harnesses whose gains transfer to unseen production harnesses; environment hardening against reward hacking; and MOPD2 multi-prefix multi-teacher on-policy distillation after RL. Pro is a 1.02T-total / 42B-active sparse MoE with 1M-token omnimodal context.

Keywords

MiMo-V2.6 · You Only RL Once · GRPO · groupwise agentic grading · GRS · GAR · MOPD2 · mini-harnesses · reward hacking · DeepSWE · Xiaomi

Topics

reinforcement learning, GRPO, agentic RL, multi-harness training, reward hacking, on-policy distillation, MoE

Research notes

  • Discovery: @SergioPaniego X thread 2026-10-01 (https://x.com/SergioPaniego/status/2105666293697527994) reading the post-training sections
  • Tech report PDF: https://huggingface.co/XiaomiMiMo/MiMo-V2.6-Pro-RL/blob/main/MiMo_V2_6_technical_report.pdf
  • Scale of the run: 30 RL steps, ~750K trajectories in under 6 days (123h Pro / 82h Flash), reported cost ~$2.62M (Pro) and ~$0.85M (Flash); per step 1,568 prompts x 16 rollouts = ~25K trajectories, 2.7-3.7B tokens per update; Pro budget split ~43.5% training / 43.8% rollouts / 12.7% grading; code is 68% of the mix yet gains show on general agent benchmarks; average training-task pass rates rose ~12% relative (Pro) and ~25% (Flash)
  • Minimal harnesses: production harnesses ship prompts/safeguards the reward never measures, so the model learns to ignore them; training on a family of minimal mini-harnesses (open source: github.com/XiaomiMiMo/mimoagent) transfers to harnesses never seen in training (Codex, Claude Code, mini-swe-agent; DeepSWE ~50% -> 66% held-out)
  • Ranking passing solutions: if all 16 rollouts pass, advantage is zero; GRS builds task-specific rubrics offline and GAR ranks passing trajectories online — without it the policy drifted into exception swallowing and relaxed validation; same ranking applied to visual-task renders
  • Reward hacking: report quotes the model fetching the upstream fix from a GitHub raw URL; countermeasures were env cleanup, network isolation, and a dedicated hack agent attacking every env; confirmed hacks stayed under 2% of the run
  • MoE detail: with a trainable router, 22% of experts went cold within 20 steps, so the router was frozen during RL
  • MOPD2 after RL: domain-specialized teachers distilled back on-policy; a k-turn trajectory gives k prefix starting points and the student generates only the next turn from each
  • Final evals (report table): DeepSWE v1.1 Pro 71.9 / Flash 67.9, AutomationBench v1.0.6 53.1 / 52.3, CyberGym 94.0 / 95.1, OSWorld-Verified 82.0 / 80.8
  • Architecture: sparse MoE 1.02T total / 42B activated, 384 routed experts (8 active), 1M context, native omnimodal (text/image/video/audio) with 681M MiMo ViT and audio encoders, 5-layer MTP speculative decoder
  • Releases: open weights Pro-RL and Flash-RL (Hugging Face + ModelScope, collection https://huggingface.co/collections/XiaomiMiMo/mimo-v26-6ab18b0eff1d27b901d9f96b), 7.8K curated RL envs (catalog Datasets row 72 MiMo-V2.6-RL-oss; FineEnvs Harbor ports rows 74-79; Explorer Space https://huggingface.co/spaces/FineEnvs/MiMo-RL-Envs-Explorer), and a 9B starting point (Qwen3.5-9B SFT'd on MiMo data) that GRPO on the released envs improves on all 11 reported benchmarks
  • Models released ~22 Sep 2026 per launch coverage; license not stated on the model card