MoonEP: A Perfectly Balanced Expert Parallelism Library via Dynamic Redundant Experts
- Type
- repo
- Venue
- Moonshot AI (GitHub)
- Year
- 2026
- Source
- github
- Access
- free
- Language
- en
- Added
- 2026-08-14T19:25:00Z
- Verified
- 2026-08-14T19:25:00Z
Summary
Moonshot MIT EP communication library (Kimi K3 Open Day, with FlashKDA and AgentEnv). Hard invariant: every EP rank receives exactly S×K tokens regardless of router skew. An on-GPU planner duplicates a small set of experts from the current routing, prefetches their weights, and reduces duplicated grads back to home ranks. Zero-copy permute/unpermute writes tokens straight into expert-grouped remote buffers; static S×K shapes drop per-layer host sync and stop the fragmentation that OOMs DeepEP under imbalance. On H20 EP=8 vs DeepEP v2, comm time stays nearly flat as maxvio grows while DeepEP degrades, and e2e iteration time stays flat. Training prefetch slots B=E/R; inference can use B=3–4. NVIDIA GPU; Zhenwu PPU listed as coming. Cite moonep2026.
Keywords
moonep · moe · expert-parallelism · moonshot · kimi · deepep · x
Topics
MoE training, expert parallelism, distributed systems
Research notes
- Primary: GitHub README (MIT, Python, 1079 stars / 118 forks at check). Discord/X https://x.com/Kimi_Moonshot/status/2081763086281973847 via fxtwitter. Inspired by DeepEP, Echo arXiv 2603.07685, UltraEP, AcclEP. Comm library, not a hosted corpus, so no datasets_local row.