The State-Prediction Separation Hypothesis
- Type
- paper
- Venue
- arXiv / Cornell University / Harvard
- Year
- 2026
- Source
- arxiv
- Access
- free
- Language
- en
- Added
- 2026-08-14T17:02:26Z
- Verified
- 2026-08-14T17:02:26Z
Summary
SPS inserts a dummy prediction token after every input so the input stream owns the persistent KV cache and the prediction stream emits the next token (plus a w=64 ephemeral prediction window). Pretraining 53M-1.678B on FineWeb-Edu: at 1.6B, matches a standard Transformer validation loss with 2.6x fewer tokens (pre-decay) and gains ~2-3 pp average zero-shot accuracy, with persistent KV ~1.01x Standard and decode throughput within 6-10%. 2x Memory and Delayed State ablations show extra compute/memory is not enough; gradient analysis routes future-loss onto the input stream. Code https://github.com/lil-lab/sps.
Keywords
sps · two-stream · kv-cache · state-prediction · fineweb-edu · cornell · harvard
Topics
transformer architecture, language modeling
Research notes
- Primary: arxiv abs (default nonexclusive-distrib 1.0, cs.LG/CL). Cornell LIL Lab; Brantley at Harvard. Correspondence giovanni@cs.cornell.edu. Code https://github.com/lil-lab/sps (23 stars at check); license not stated on the GitHub landing page, so left blank. Discord posted AlphaXiv 2607.01218; canonical abs recorded. HF paper page 12 upvotes, org lil-lab; no linked models/datasets. Trains on public FineWeb-Edu; no new corpus, so no datasets_local row.