← Back to explorer

ID Balancing: Stable Training of Extremely Sparse MoE via PID-Based Load Control

Type
paper
Venue
arXiv 2026-09-30 (cs.LG)
Year
2026
Source
paper
Access
public
Language
en
Added
2026-10-01
Verified
2026-10-01

Summary

Addresses expert load imbalance when scaling LLMs via Mixture-of-Experts with increasingly sparse routing. ID Balancing: an Integral-Derivative load-control method built on the observation that two existing auxiliary-loss-free methods are incomplete PID controllers -- DeepSeek's loss-free method is a fixed-step integral controller and Kimi K3's Quantile Balancing is a generalized proportional controller. ID Balancing scales its integral term with load error and activates its derivative term only when imbalance worsens, giving stronger corrections for large/worsening errors and smaller updates near balance. Results (Top-10/Top-5/Top-3 routing over 768 experts): in Top-3, worst-case backbone MaxVio reduced by over 50% and training-average backbone MinVio by over 12% vs the best baselines. Scaling 18.9B -> 69.9B total params (Top-10-of-768), worst-case backbone MaxVio stays nearly unchanged, ~89.6% lower than the auxiliary-loss baseline. Maintains competitive language-modeling and downstream performance; advantages grow as sparsity increases.

Keywords

mixture-of-experts · MoE · load balancing · PID control · sparse routing · training stability

Topics

mixture-of-experts, MoE, load balancing, PID control, sparse routing, training stability

Research notes

  • Discovery: shared directly in chat (2026-10-01)
  • No code repo, model, dataset, or demo linked on the arXiv page.
  • License: arXiv non-exclusive distribution only (no open CC license).
  • Connects to the collection's MoE/sparse-training entries and the user's own pretraining work.