← Back to explorer

Gated DeltaNet-2: Decoupling Erase and Write in Linear Attention

Type
paper
Venue
arXiv:2605.22791
Year
2026
Source
arxiv
Access
public
Language
en
Added
2026-09-29
Verified
2026-09-29

Summary

Gated DeltaNet-2 (NVIDIA) is a linear-attention architecture that splits the delta rule's single scalar gate into independent channel-wise erase and write gates, letting the model forget old associations and commit new ones without interference. It achieves the best results among Mamba-2/3 and DeltaNet-family variants at 1.3B scale while keeping linear-time decoding, though finite state capacity still limits multi-needle recall.

Keywords

linear-attention · ssm · architecture · long-context · nvidia

Topics

linear-attention, ssm, architecture, long-context, nvidia

Research notes

  • Method: Replace the scalar gate of Gated DeltaNet / Kimi Delta Attention with two independent channel-wise gates — erase on the key-side read direction, write on the value-side commit direction — plus channel-wise decay; derive a chunkwise parallel training algorithm absorbing channel-wise decay into asymmetric rank-one erase factors.
  • Key findings: Decouples the single scalar delta gate into a channel-wise erase gate (key side) and channel-wise write gate (value side): S_t = Diag(α_t)(I − b_t k_t k_tᵀ)S_{t−1} + w_t k_t v_tᵀ; Strictly generalizes predecessors: collapsing both gates to one scalar recovers Kimi Delta Attention; collapsing decay too recovers Gated DeltaNet; At 1.3B params / 100B FineWeb-Edu tokens: strongest among Mamba-2/3, Gated DeltaNet, KDA on language modeling, commonsense reasoning, and long-context retrieval; largest edge on RULER needle-in-a-haystack; Chunkwise WY formulation + Triton kernels preserve efficient parallel training with a gate-aware backward pass
  • Limitations: Purely linear layers remain finite-state-capacity limited on multi-item associative recall (RULER multi-needle exposes this); the architecture retains full-attention layers rather than removing them.
  • NVIDIA (NVlabs). Submitted May 21, 2026. Official PyTorch implementation released.