← Back to explorer

LongCat Sparse Attention: Taming the Lightning via Streaming-aware Hierarchical Cross-Layer Indexing

Type
paper
Venue
arXiv / Meituan LongCat
Year
2026
Source
arxiv
Access
free
Language
en
Added
2026-08-14T16:24:56Z
Verified
2026-08-14T16:24:56Z

Summary

LongCat Sparse Attention (LSA) composes three indexer changes on DSA: Streaming-Aware Indexing (sink+sliding-window contiguous budget plus dynamic sparse tokens for coalesced HBM), Cross-Layer Indexing (N=2 owner/reuse with distillation; halves indexer passes), and training-free Hierarchical Indexing (coarse page recall then token refine; net win at >=256K). Matches full MLA on general/long-context benches at 69B-A3B and 560B-A27B, enables 1M-token native training, and underpins LongCat-2.0 (1.6T-A48B). Releases LongCat-Flash-Lite-Sparse (69B-A3B).

Keywords

sparse-attention · dsa · lightning-indexer · long-context · longcat · meituan · moe · mla

Topics

sparse attention, long-context LLMs

Research notes

  • Primary: arxiv abs (default nonexclusive-distrib 1.0; cs.AI/CL/DC/LG). Corresponding zhangjiaqi39@meituan.com. Open model https://huggingface.co/meituan-longcat/LongCat-Flash-Lite-Sparse (MIT, 69B, 82 likes, 1.1k downloads). LongCat-2.0 also MIT. Discord posted AlphaXiv 2608.01662v1. Paper not a dataset; no datasets_local row.