Diagnosing and Mitigating Tool-Call Repetition in MiMo-V2.6
- Type
- blog
- Venue
- MiMo blog (mimo.xiaomi.com)
- Year
- 2026
- Source
- web
- Access
- public
- Language
- en
- Added
- 2026-09-29
- Verified
- 2026-09-29
Summary
Post-mortem on MiMo-V2.6 models repeating identical or near-identical tool calls in agentic settings (MiMo Desktop, MiMo Code, OpenCode), burning context and stalling tasks. Root cause: a reward blind spot — the RL flooding penalty only triggered above 32 tool calls per turn, so sub-threshold high-volume calling was amplified during training. Fix: a repetition-specialized single-turn RL teacher merged into the main model via MOPD (Multi-teacher On-Policy Distillation) with teacher-prefix OPD.
Keywords
RL · agentic · tool-use · post-mortem · distillation · MiMo
Topics
RL, agentic, tool-use, post-mortem, distillation
Research notes
- Also: https://x.com/XiaomiMiMoDevs/status/2104251067039191324
- Discovery: Shared in #random-papers twice on 2026-09-27: the MiMo blog post-mortem and the @XiaomiMiMoDevs X announcement of the same fix.
- Method: Reproduced repetition via exact within-turn duplication metric (same tool + JSON-canonicalized args); traced emergence across RL checkpoints (flooding share rose 11.1% at step 0 to 24.6% at step 20 for Flash-RL). Tested stricter penalty (32 to 8 calls/turn) — effective but slow (~20 steps, ~$2.31M estimated at full scale) and poor generalization (13.45% to 3.83% replay repetition). Instead trained a specialized RL teacher (reward 0 on repetition, 1 only when clean, plus KL to original policy; 12 steps, ~7,000 examples, zero replay repetition on train and held-out), then distilled it in via MOPD teacher-prefix OPD rolled back five steps.
- Key findings: The specialized teacher learned when to stop: in a 59-tool-call example trajectory, end-of-turn token probability at the 12th call position flipped from 5.43% to 92.17% (cumulative stop probability 99.87% by the 8th call after fix vs 0.32% before). After MOPD, repetition rates dropped substantially across harnesses and context lengths for both Pro and Flash while the broader benchmark suite held steady. Total cost ~$90,000 — 4% of the estimated MixRL-redo alternative.
- Limitations: Company-reported metrics on internal evaluation sets; the exact within-turn metric is a lower bound that excludes cross-turn repetition and near-duplicates.
- Details reference Section 5.6 of the MiMo-V2.6 technical report (Hugging Face). X post dated 2026-09-27 12:45 PM; blog published 2026-09-27.