xattn
- Type
- repo
- Venue
- GitHub
- Year
- 2026
- Source
- github
- Access
- public
- Language
- en
- Added
- 2026-09-29
- Verified
- 2026-09-29
Summary
Efficient attention modules for the xLLM framework, with Python APIs for use in other PyTorch projects. Features high-precision FP32 attention output state for backward (reducing BF16 gradient error) and flexible segment-aware masking (segment IDs or BOS masks) for packed sequences. Supports causal Flash Attention, sliding window, sliding chunk, and SoftDelta attention variants; FP16/BF16 inputs; MHA, GQA, and MQA.
Keywords
attention · FlashAttention · CUDA kernels · long context · PyTorch
Topics
attention kernels, LLM training infrastructure
Research notes
- Discovery: open-sourced by Xuezhe "Max" Ma (@MaxMa1987, Research Lead @ USC ISI / Asst. Prof. @ USC CS; PhD CMU) on 2026-09-28, quoting an @IFM_AI post on LLM pre-training/fine-tuning rework costs: https://x.com/MaxMa1987/status/2104650088610156753
- Method: CUDA attention kernels.
- Apache 2.0.
- Created 2026-09-25; 6 stars at research time.
- Led by Shicheng Wen (per announcement thread).
- SM90 builds supported; SM100+ not yet supported.