← Back to explorer

MoBA: Mixture of Block Attention for Long-Context LLMs

Type
paper
Venue
arXiv (cs.LG)
Year
2026
Source
arxiv
Access
public
Language
en
Added
2026-09-29
Verified
2026-09-29

Summary

Proposes Mixture of Block Attention (MoBA), which applies mixture-of-experts principles to the attention mechanism: the model learns where to attend across blocks rather than using predefined sparse patterns like sink or window attention. It claims superior long-context performance with seamless switching between full and sparse attention, and has been deployed for Kimi's long-context serving.

Keywords

attention · long context · mixture of experts · efficient inference · Kimi

Topics

attention, long context, mixture of experts, efficient inference, Kimi

Research notes

  • Method: Block-wise sparse attention with learned MoE-style routing over blocks, following a 'less structure' principle so the model determines attention targets autonomously instead of using predefined sparse patterns.
  • Key findings: Deployed to support Kimi's long-context requests; Claims superior long-context performance with seamless switching between full and sparse attention
  • Limitations: License CC-BY-NC-ND 4.0; code link on the arXiv page was generic/unverified.
  • arXiv:2502.13189, cs.LG, submitted 2025-02-18.