Parallax: Parameterized Local Linear Attention for Language Modeling
- Type
- other
- Venue
- arXiv / Northwestern University / Tilde Research / University of Washington
Summary
Parallax parameterizes LLA's local-linear correction as ρ=W_R x, dropping the CG solve and the unstable boundary-amplification factor. A FlashAttention-style streaming kernel reuses the KV stream and raises arithmetic intensity; a Hopper decode prototype matches or beats FA2/3. Pretrained at 0.6B (78.6B tokens) and 1.7B (157.2B) on Ultra-FineWeb with a Qwen-3 backbone: under Muon, Parallax beats softmax on LAMBADA/WikiText and zero-shot avg (0.6B 55.99 vs Transformer 54.54; 1.7B 62.45 vs 61.43). Gains hold under parameter-matched and compute-matched controls. Under AdamW the correction branch collapses and the gap vanishes — first reported architecture–optimizer codesign for attention. Code https://github.com/Yifei-Zuo/Parallax.
Keywords
parallax · lla · muon · flashattention · northwestern · tilde · test-time-regression
Topics
linear attention, local linear attention, LLM pretraining
Research notes
- Primary: arxiv abs (cs.LG/AI/CL). Northwestern / Tilde Research / UW; correspondence yifeizuo2029@u.northwestern.edu, dhruv@tilderesearch.com, zhaoranwang@gmail.com. Code MIT https://github.com/Yifei-Zuo/Parallax (68 stars at check); Discord also posted this GitHub in 1510318176196628682. HF paper page 11 upvotes, org Northwestern University. Linked models YifeiZuo/Parallax-0.6B (655M, 11 dl) and Parallax-1.7B (1.84B, 9 dl). Discord posted abs. Pretrains on public Ultra-FineWeb; no new corpus, so no datasets_local row. ArXiv license widget not visible in converted abs HTML, so paper license left blank.