← Back to explorer

xattn

Type
repo
Venue
GitHub
Year
2026
Source
github
Access
public
Language
en
Added
2026-09-29
Verified
2026-09-29

Summary

Efficient attention modules for the xLLM framework, with Python APIs for use in other PyTorch projects. Features high-precision FP32 attention output state for backward (reducing BF16 gradient error) and flexible segment-aware masking (segment IDs or BOS masks) for packed sequences. Supports causal Flash Attention, sliding window, sliding chunk, and SoftDelta attention variants; FP16/BF16 inputs; MHA, GQA, and MQA.

Keywords

attention · FlashAttention · CUDA kernels · long context · PyTorch

Topics

attention kernels, LLM training infrastructure

Research notes

  • Discovery: open-sourced by Xuezhe "Max" Ma (@MaxMa1987, Research Lead @ USC ISI / Asst. Prof. @ USC CS; PhD CMU) on 2026-09-28, quoting an @IFM_AI post on LLM pre-training/fine-tuning rework costs: https://x.com/MaxMa1987/status/2104650088610156753
  • Method: CUDA attention kernels.
  • Apache 2.0.
  • Created 2026-09-25; 6 stars at research time.
  • Led by Shicheng Wen (per announcement thread).
  • SM90 builds supported; SM100+ not yet supported.