Gated DeltaNet-2: Decoupling Erase and Write in Linear Attention
- Type
- paper
- Venue
- arXiv:2605.22791
- Year
- 2026
- Source
- arxiv
- Access
- public
- Language
- en
- Added
- 2026-09-29
- Verified
- 2026-09-29
Summary
Gated DeltaNet-2 (NVIDIA) is a linear-attention architecture that splits the delta rule's single scalar gate into independent channel-wise erase and write gates, letting the model forget old associations and commit new ones without interference. It achieves the best results among Mamba-2/3 and DeltaNet-family variants at 1.3B scale while keeping linear-time decoding, though finite state capacity still limits multi-needle recall.
Keywords
linear-attention · ssm · architecture · long-context · nvidia
Topics
linear-attention, ssm, architecture, long-context, nvidia
Research notes
- Method: Replace the scalar gate of Gated DeltaNet / Kimi Delta Attention with two independent channel-wise gates — erase on the key-side read direction, write on the value-side commit direction — plus channel-wise decay; derive a chunkwise parallel training algorithm absorbing channel-wise decay into asymmetric rank-one erase factors.
- Key findings: Decouples the single scalar delta gate into a channel-wise erase gate (key side) and channel-wise write gate (value side): S_t = Diag(α_t)(I − b_t k_t k_tᵀ)S_{t−1} + w_t k_t v_tᵀ; Strictly generalizes predecessors: collapsing both gates to one scalar recovers Kimi Delta Attention; collapsing decay too recovers Gated DeltaNet; At 1.3B params / 100B FineWeb-Edu tokens: strongest among Mamba-2/3, Gated DeltaNet, KDA on language modeling, commonsense reasoning, and long-context retrieval; largest edge on RULER needle-in-a-haystack; Chunkwise WY formulation + Triton kernels preserve efficient parallel training with a gate-aware backward pass
- Limitations: Purely linear layers remain finite-state-capacity limited on multi-item associative recall (RULER multi-needle exposes this); the architecture retains full-attention layers rather than removing them.
- NVIDIA (NVlabs). Submitted May 21, 2026. Official PyTorch implementation released.