← Back to explorer

FlashAttention-3: Fast and Accurate Attention with Asynchrony and Low-precision

Type
other
Venue
arXiv / Colfax International / Meta / NVIDIA

Summary

FlashAttention-3 redesigns FA2 for Hopper: warp-specialized TMA/WGMMA producer-consumer, pingpong scheduling so softmax hides under GEMM, and FP8 with in-kernel V transpose, block quantization, and Hadamard incoherent processing. H100 FP16 is 1.5-2.0× FA2 (up to 740 TFLOPs/s, 75% util) and 1.5-1.75× backward; FP8 reaches ~1.2 PFLOPs/s with 2.6× lower RMSE vs per-tensor FP8 on outlier-injected QKV. Competitive with or faster than cuDNN FA2 at seq≥1k. Code https://github.com/Dao-AILab/flash-attention.

Keywords

flashattention-3 · hopper · wgmma · tma · fp8 · warp-specialization · colfax · nvidia · meta · attention

Topics

attention kernels, Hopper GPUs, FP8

Research notes

  • Primary: arxiv abs (cs.LG; also cs.AI). License CC BY 4.0 on HTML at check. Equal contrib Shah/Bikshandi (Colfax); Zhang Meta; Thakkar/Ramani NVIDIA; Dao tri@tridao.me. Code https://github.com/Dao-AILab/flash-attention (24,706 stars at check; BSD-3-Clause). HF paper page 2 upvotes; unofficial linked models ethanker/nanomind-step-002000 and AshenNav/twill-swp-ws not copied into hf_* fields. Discord posted abs. Kernel library, not a new corpus, so no datasets_local row.