← Back to explorer

Speculative decoding with Gated DeltaNet: no un-reading needed

Type
x_thread
Venue
X
Year
2026
Source
x
Access
public
Language
en
Added
2026-09-29
Verified
2026-09-29

Summary

Standalone technical insight on whether speculative decoding is a problem for Gated DeltaNet. Argument: rejected drafts must be rolled back; attention just drops their KV rows, but GDN has already blended every draft into its recurrent state. However, the update can be rewritten so that only accepted drafts are committed: S_t = alpha_t S_{t-1} + v_tilde_t k_t^T, with v_tilde_t = beta_t (v_t - alpha_t S_{t-1} k_t), giving S_a = g_a S_0 + sum_{i<=a} (g_a/g_i) v_tilde_i k_i^T (g = product of alphas). Cache (v_tilde_i, k_i, g_i) per draft and fold in only what is accepted. Conclusion: GDN does not have to un-read a token; it can wait to commit it.

Keywords

speculative decoding · Gated DeltaNet · linear attention · KV cache

Topics

speculative decoding, linear attention, recurrent models

Research notes

  • Discovery: link pasted by user; tweet by Ali Hatamizadeh (@ahatamiz1, LLM Tech Lead & Staff Research Scientist @ NVIDIA, co-creator of Gated DeltaNet / Gated DeltaNet-2), posted 2026-09-29 3:24 PM. No linked paper/repo/dataset - the tweet itself is the artifact. A reply by @liranringel asked whether draft trees are an issue (naive implementations would need extra state per branch).