← Back to explorer

Gated DeltaNet-2: Decoupling Erase and Write in Linear Attention

Type
paper
Venue
arXiv / NVIDIA
Year
2026
Source
arxiv
Access
free
Language
en
Added
2026-08-14T19:20:00Z
Verified
2026-08-14T19:20:00Z

Summary

NVIDIA linear-attention layer that generalizes Gated DeltaNet and Kimi Delta Attention by replacing the tied scalar delta gate with a key-side erase gate b_t and a value-side write gate w_t, keeping channel-wise decay. Recovers KDA when both gates collapse to the same scalar and Gated DeltaNet when decay is scalar too. Derives a fast-weight view, chunkwise WY algorithm, and gate-aware backward. At 1.3B on 100B FineWeb-Edu tokens, strongest overall vs Mamba-2/GDN/KDA/Mamba-3 on LM, commonsense, and retrieval; largest gains on RULER multi-key NIAH. Recurrent real-world retrieval avg 29.88 vs KDA 28.67; hybrid 42.28 vs 40.14. Code NVlabs/GatedDeltaNet-2.

Keywords

gated-deltanet · kda · mamba · linear-attention · nvidia · delta-rule

Topics

linear attention, delta rule, recurrent LMs

Research notes

  • Primary: arxiv abs 2605.22791. Discord posted https://t.co/Zw6yXbHjGU which resolves to https://github.com/NVlabs/GatedDeltaNet-2/blob/main/paper/GDN2_paper.pdf. Code GitHub SPDX Other / NVIDIA Source Code License-NC per README coverage; license field left blank rather than guess SPDX. Architecture/code item trained on public FineWeb-Edu, not a new hosted corpus, so no datasets_local row.