MoBA: Mixture of Block Attention for Long-Context LLMs
- Type
- paper
- Venue
- arXiv (cs.LG)
- Year
- 2026
- Source
- arxiv
- Access
- public
- Language
- en
- Added
- 2026-09-29
- Verified
- 2026-09-29
Summary
Proposes Mixture of Block Attention (MoBA), which applies mixture-of-experts principles to the attention mechanism: the model learns where to attend across blocks rather than using predefined sparse patterns like sink or window attention. It claims superior long-context performance with seamless switching between full and sparse attention, and has been deployed for Kimi's long-context serving.
Keywords
attention · long context · mixture of experts · efficient inference · Kimi
Topics
attention, long context, mixture of experts, efficient inference, Kimi
Research notes
- Method: Block-wise sparse attention with learned MoE-style routing over blocks, following a 'less structure' principle so the model determines attention targets autonomously instead of using predefined sparse patterns.
- Key findings: Deployed to support Kimi's long-context requests; Claims superior long-context performance with seamless switching between full and sparse attention
- Limitations: License CC-BY-NC-ND 4.0; code link on the arXiv page was generic/unverified.
- arXiv:2502.13189, cs.LG, submitted 2025-02-18.