Flash-MSA: Accelerating Million-Token Training With Sparse Attention Kernels
- Type
- other
- Venue
- nanduruganesh.github.io
- Year
- 2026
- Source
- web
- Access
- free
- Language
- en
- Added
- 2026-08-14T19:06:00Z
- Verified
- 2026-08-14T19:06:00Z
Summary
Flash-MSA implements MSA training (proxy block-sparse attention + fused proxy/main backward with an atomic KL trick so proxy grad = p_proxy − p_main) in CuTeDSL for H100/B200, CUDA 13. Tweet claims 400%+ speedup vs dense flash at long context. Warmup kernels run dense flash then train the proxy in the backward. Correctness vs eager PyTorch cosine sim typically ≥0.998. Requires adding kernel KL loss to CE. Inference kernels remain MiniMax-AI/MSA.
Keywords
flash-msa · sparse-attention · cute · hopper · blackwell · minimax · kernels · blog · x
Topics
sparse attention, CUDA kernels, long-context training
Research notes
- Primary: technical blog. Discord/X https://x.com/gnanduru1/status/2076342414915342718 via fxtwitter. Code https://github.com/nanduruganesh/flash-msa (MIT); PyPI flash-msa; Megatron-LM fork https://github.com/nanduruganesh/Megatron-LM. Underlying MSA paper arXiv 2606.13392; official inference https://github.com/MiniMax-AI/MSA. Kernel/code item, not a hosted corpus, so no datasets_local row.