Screening Is Enough
- Type
- other
- Venue
- arXiv / RIKEN CEMS / The University of Tokyo
Summary
Multiscreen screens unit-normalized QK similarities with Trim (exact zero below threshold) and a learned causal Softmask window, then aggregates surviving values without competition; MiPE rotates only two dims and turns off for large windows. On SlimPajama (2^38 tokens, seq 2^12) it matches Transformer val loss with ~30% fewer params, stays stable at LR 2^{-4} (Transformer diverges), holds long-context PPL past train length, and on ABCDigits retrieval a 286M Multiscreen beats a 1.3B Transformer (99.18% vs 95.98% at train ctx; little degradation to 2^17). Full-context forward latency lower than FlashAttention Transformer at long ctx (RTX 4090, no KV cache). ABCDigits generator https://github.com/ken-nakanishi/abcdigits.
Keywords
multiscreen · screening · long-context · attention · abcdigits · rope · mamba-alternative · riken
Topics
attention, long context, architecture
Research notes
- Primary: arxiv abs (cs.LG; also cs.AI, cs.CL). License not stated on abs/HTML at check. Ken M. Nakanishi, RIKEN CEMS / UTokyo (ken.m.nakanishi@gmail.com). ABCDigits code https://github.com/ken-nakanishi/abcdigits (GitHub API rate-limited at check). No official model code on abs. HF paper page 0 upvotes; no linked models/datasets. Discord posted abs. ABCDigits is a prompt generator, not a hosted corpus, so no datasets_local row. License field left blank per catalog convention.