← Back to explorer

SageAttention: Accurate 8-Bit Attention for Plug-and-play Inference Acceleration

Type
other
Venue
arXiv / Tsinghua University

Summary

PTQ attention: subtract the token-mean of K to smooth channel outliers so QK can be INT8, and keep PV in FP16 with an FP16 accumulator (INT8 PV is too noisy on some layers). Triton kernels on RTX4090/3090; peak ~341 TOPS at headdim 64/128 vs FlashAttention2 ~165. ~2.1× FA2 and ~2.7× xformers. Plug-and-play: Llama2-7B WikiText 5.824 vs FP 5.823, MMLU 0.46=FP; CogVideoX/Unidiffuser/UltraPixel/TIMM/Llava1.6 near-parity. ICLR 2025. Code https://github.com/thu-ml/SageAttention.

Keywords

sageattention · int8 · flashattention · quantization · triton · iclr · tsinghua · inference

Topics

quantization, attention kernels, inference

Research notes

  • Primary: arxiv abs (cs.LG). License CC BY 4.0 on HTML at check. Comment: ICLR 2025 bibitem; journal_ref The Thirteenth International Conference on Learning Representations (ICLR 2025). Corresponding Jianfei Chen; Tsinghua CST / Institute for AI / BNRist / Tsinghua-Bosch Joint ML Center. Code https://github.com/thu-ml/SageAttention (HF projectPage; 3,640 stars at check; Apache-2.0). HF paper page 50 upvotes; unofficial linked model jt-zhang/SageAttention2_plus not copied into hf_* fields. Discord posted abs with a speed/accuracy blurb. Kernel library, not a new corpus, so no datasets_local row.