Full-Pipeline Inference Optimization for MiMo-V2.5 Series: Pushing Hybrid SWA Efficiency to the Limit
- Type
- paper
- Venue
- arXiv / Xiaomi MiMo
- Year
- 2026
- Source
- arxiv
- Access
- free
- Language
- en
- Added
- 2026-08-14T16:40:00Z
- Verified
- 2026-08-14T16:40:00Z
Summary
Engineering report for MiMo-V2.5 / V2.5-Pro Hybrid Sliding Window Attention (example: 70 layers, 10 full + 60 SWA, W=128; ~7× less attention FLOPs and KV vs full attention) plus sparse MoE and vision/audio/video encoders. Dual-pool KV with strict O(W) SWA storage, layerwise prefetch, SWA-aware prefix cache trees, GCache L3 (RDMA, co-deployed on GPU nodes), KVCache-affinity router (~+25% L2 hit, ~+30% input throughput), length bucketing, halved EP after SWA storage (~40% e2e), MTP in prefill (2.3× on first 128 decode tokens), GPU image preprocess and parallel video decode (1-hour video 156s→23s). Reports ~93% production KV hit rate. Positions as the first large-scale serving system covering Hybrid SWA+MoE+multimodal. Not a dataset.
Keywords
mimo · hybrid-swa · kv-cache · sglang · moe · serving · xiaomi · multimodal · gcache
Topics
LLM serving, Hybrid SWA, KV cache, multimodal inference
Research notes
- Primary: arxiv abs (default nonexclusive-distrib 1.0, cs.AR/AI); technical report. Xiaomi MiMo Team; corresponding Fuli Luo. Open weights https://huggingface.co/XiaomiMiMo/MiMo-V2.5 (MIT, 394 likes, 428,880 downloads at check) and MiMo-V2.5-Pro. Serving work on SGLang; some Encoder changes noted as SGLang issue #24945. Discord posted AlphaXiv 2607.13095 with a chatId query; canonical abs recorded. Paper not a dataset; no datasets_local row.