← Back to explorer

Diagnosing and Mitigating Tool-Call Repetition in MiMo-V2.6

Type
blog
Venue
MiMo blog (mimo.xiaomi.com)
Year
2026
Source
web
Access
public
Language
en
Added
2026-09-29
Verified
2026-09-29

Summary

Post-mortem on MiMo-V2.6 models repeating identical or near-identical tool calls in agentic settings (MiMo Desktop, MiMo Code, OpenCode), burning context and stalling tasks. Root cause: a reward blind spot — the RL flooding penalty only triggered above 32 tool calls per turn, so sub-threshold high-volume calling was amplified during training. Fix: a repetition-specialized single-turn RL teacher merged into the main model via MOPD (Multi-teacher On-Policy Distillation) with teacher-prefix OPD.

Keywords

RL · agentic · tool-use · post-mortem · distillation · MiMo

Topics

RL, agentic, tool-use, post-mortem, distillation

Research notes

  • Also: https://x.com/XiaomiMiMoDevs/status/2104251067039191324
  • Discovery: Shared in #random-papers twice on 2026-09-27: the MiMo blog post-mortem and the @XiaomiMiMoDevs X announcement of the same fix.
  • Method: Reproduced repetition via exact within-turn duplication metric (same tool + JSON-canonicalized args); traced emergence across RL checkpoints (flooding share rose 11.1% at step 0 to 24.6% at step 20 for Flash-RL). Tested stricter penalty (32 to 8 calls/turn) — effective but slow (~20 steps, ~$2.31M estimated at full scale) and poor generalization (13.45% to 3.83% replay repetition). Instead trained a specialized RL teacher (reward 0 on repetition, 1 only when clean, plus KL to original policy; 12 steps, ~7,000 examples, zero replay repetition on train and held-out), then distilled it in via MOPD teacher-prefix OPD rolled back five steps.
  • Key findings: The specialized teacher learned when to stop: in a 59-tool-call example trajectory, end-of-turn token probability at the 12th call position flipped from 5.43% to 92.17% (cumulative stop probability 99.87% by the 8th call after fix vs 0.32% before). After MOPD, repetition rates dropped substantially across harnesses and context lengths for both Pro and Flash while the broader benchmark suite held steady. Total cost ~$90,000 — 4% of the estimated MixRL-redo alternative.
  • Limitations: Company-reported metrics on internal evaluation sets; the exact within-turn metric is a lower bound that excludes cross-turn repetition and near-duplicates.
  • Details reference Section 5.6 of the MiMo-V2.6 technical report (Hugging Face). X post dated 2026-09-27 12:45 PM; blog published 2026-09-27.