Dynamic Linear Attention
- Type
- paper
- Venue
- arXiv / Ohio State / University of Michigan / ByteDance Seed
- Year
- 2026
- Source
- arxiv
- Access
- free
- Language
- en
- Added
- 2026-08-14T17:02:26Z
- Verified
- 2026-08-14T17:02:26Z
Summary
DLA builds multi-state linear-attention memory with (i) Information-Aware Dynamic State Merging: a State Information Score opens a new state when token-level Frobenius drift exceeds tau, else merges; (ii) Capacity-Bounded Memory: a K-slot chronological cache that merges the lowest-density adjacent pair when full. Pretrains Mamba-2-780M and Gated DeltaNet-1.3B on 50B Long-Data-Collections tokens (seq 16k, K=30). Beats Log-Linear Attention on 8 commonsense, 6 in-context retrieval, RULER, and LongBench; Mamba-2+DLA matches or beats a 778M Transformer. Higher prefill throughput and lower memory than log-linear. No official code on the abs page.
Keywords
dla · linear-attention · mamba-2 · gated-deltanet · long-context · bytedance · osu
Topics
linear attention, long context
Research notes
- Primary: arxiv abs (default nonexclusive-distrib 1.0, cs.LG/CL). Ohio State / Michigan / ByteDance Seed; correspondence wang.15980@osu.edu, mizhang.1@osu.edu. Discord posted AlphaXiv replicate 2606.10650; canonical abs recorded. HF paper page 10 upvotes, org ByteDance-Seed; no linked models/datasets. No official code on abs. Pretrains on existing Long-Data-Collections; no new corpus, so no datasets_local row.