Gated Attention for Large Language Models: Non-linearity, Sparsity, and Attention-Sink-Free
- Type
- other
- Venue
- arXiv / Qwen Team / Alibaba
Summary
Compares ~30 gating variants on 15B MoE (15A2B, 2.54B active) and 1.7B dense models trained on up to 3.5T tokens. A head-specific sigmoid gate after SDPA (G1) is best: MoE 400B-token Avg PPL 5.761 vs 6.026 baseline and MMLU 60.82 vs 58.79, beating parameter-matched KV-head/expert expansions. Gains come from non-linearity between low-rank Wv and Wo plus query-dependent sparse gates (mean score 0.116). First-token attention falls 46.7%→4.8%; YaRN-extended RULER at 128k is 58.82 vs 31.65. Gate also damps loss spikes and allows higher LR (8e-3) where the baseline diverges. Claims first attention-sink-free models.
Keywords
gated-attention · attention-sink · sparsity · moe · qwen · alibaba · ruler
Topics
attention, architecture, LLM pretraining
Research notes
- Primary: arxiv abs (cs.CL). License not stated on abs/HTML at check. Qwen Team/Alibaba with Edinburgh (Huang), Stanford (Wen), MIT (Yang), Tsinghua (S. Huang). Code https://github.com/qiuzh20/gated_attention (973 stars at check). HF paper page 11 upvotes; 10 linked models (unofficial) not copied into hf_* fields. Discord posted abs. Trains on an internal 3.5T mix; no new public corpus, so no datasets_local row. License field left blank per catalog convention.