ID Balancing: Stable Training of Extremely Sparse MoE via PID-Based Load Control
- Type
- paper
- Venue
- arXiv 2026-09-30 (cs.LG)
- Year
- 2026
- Source
- paper
- Access
- public
- Language
- en
- Added
- 2026-10-01
- Verified
- 2026-10-01
Summary
Addresses expert load imbalance when scaling LLMs via Mixture-of-Experts with increasingly sparse routing. ID Balancing: an Integral-Derivative load-control method built on the observation that two existing auxiliary-loss-free methods are incomplete PID controllers -- DeepSeek's loss-free method is a fixed-step integral controller and Kimi K3's Quantile Balancing is a generalized proportional controller. ID Balancing scales its integral term with load error and activates its derivative term only when imbalance worsens, giving stronger corrections for large/worsening errors and smaller updates near balance. Results (Top-10/Top-5/Top-3 routing over 768 experts): in Top-3, worst-case backbone MaxVio reduced by over 50% and training-average backbone MinVio by over 12% vs the best baselines. Scaling 18.9B -> 69.9B total params (Top-10-of-768), worst-case backbone MaxVio stays nearly unchanged, ~89.6% lower than the auxiliary-loss baseline. Maintains competitive language-modeling and downstream performance; advantages grow as sparsity increases.
Keywords
mixture-of-experts · MoE · load balancing · PID control · sparse routing · training stability
Topics
mixture-of-experts, MoE, load balancing, PID control, sparse routing, training stability
Research notes
- Discovery: shared directly in chat (2026-10-01)
- No code repo, model, dataset, or demo linked on the arXiv page.
- License: arXiv non-exclusive distribution only (no open CC license).
- Connects to the collection's MoE/sparse-training entries and the user's own pretraining work.