Gated DeltaNet-2: Decoupling Erase and Write in Linear Attention
- Type
- paper
- Venue
- arXiv / NVIDIA
- Year
- 2026
- Source
- arxiv
- Access
- free
- Language
- en
- Added
- 2026-08-14T19:20:00Z
- Verified
- 2026-08-14T19:20:00Z
Summary
NVIDIA linear-attention layer that generalizes Gated DeltaNet and Kimi Delta Attention by replacing the tied scalar delta gate with a key-side erase gate b_t and a value-side write gate w_t, keeping channel-wise decay. Recovers KDA when both gates collapse to the same scalar and Gated DeltaNet when decay is scalar too. Derives a fast-weight view, chunkwise WY algorithm, and gate-aware backward. At 1.3B on 100B FineWeb-Edu tokens, strongest overall vs Mamba-2/GDN/KDA/Mamba-3 on LM, commonsense, and retrieval; largest gains on RULER multi-key NIAH. Recurrent real-world retrieval avg 29.88 vs KDA 28.67; hybrid 42.28 vs 40.14. Code NVlabs/GatedDeltaNet-2.
Keywords
gated-deltanet · kda · mamba · linear-attention · nvidia · delta-rule
Topics
linear attention, delta rule, recurrent LMs
Research notes
- Primary: arxiv abs 2605.22791. Discord posted https://t.co/Zw6yXbHjGU which resolves to https://github.com/NVlabs/GatedDeltaNet-2/blob/main/paper/GDN2_paper.pdf. Code GitHub SPDX Other / NVIDIA Source Code License-NC per README coverage; license field left blank rather than guess SPDX. Architecture/code item trained on public FineWeb-Edu, not a new hosted corpus, so no datasets_local row.