← Back to explorer

Flash-MSA: Accelerating Million-Token Training With Sparse Attention Kernels

Type
other
Venue
nanduruganesh.github.io
Year
2026
Source
web
Access
free
Language
en
Added
2026-08-14T19:06:00Z
Verified
2026-08-14T19:06:00Z

Summary

Flash-MSA implements MSA training (proxy block-sparse attention + fused proxy/main backward with an atomic KL trick so proxy grad = p_proxy − p_main) in CuTeDSL for H100/B200, CUDA 13. Tweet claims 400%+ speedup vs dense flash at long context. Warmup kernels run dense flash then train the proxy in the backward. Correctness vs eager PyTorch cosine sim typically ≥0.998. Requires adding kernel KL loss to CE. Inference kernels remain MiniMax-AI/MSA.

Keywords

flash-msa · sparse-attention · cute · hopper · blackwell · minimax · kernels · blog · x

Topics

sparse attention, CUDA kernels, long-context training

Research notes

  • Primary: technical blog. Discord/X https://x.com/gnanduru1/status/2076342414915342718 via fxtwitter. Code https://github.com/nanduruganesh/flash-msa (MIT); PyPI flash-msa; Megatron-LM fork https://github.com/nanduruganesh/Megatron-LM. Underlying MSA paper arXiv 2606.13392; official inference https://github.com/MiniMax-AI/MSA. Kernel/code item, not a hosted corpus, so no datasets_local row.