← Back to explorer

The State-Prediction Separation Hypothesis

Type
paper
Venue
arXiv / Cornell University / Harvard
Year
2026
Source
arxiv
Access
free
Language
en
Added
2026-08-14T17:02:26Z
Verified
2026-08-14T17:02:26Z

Summary

SPS inserts a dummy prediction token after every input so the input stream owns the persistent KV cache and the prediction stream emits the next token (plus a w=64 ephemeral prediction window). Pretraining 53M-1.678B on FineWeb-Edu: at 1.6B, matches a standard Transformer validation loss with 2.6x fewer tokens (pre-decay) and gains ~2-3 pp average zero-shot accuracy, with persistent KV ~1.01x Standard and decode throughput within 6-10%. 2x Memory and Delayed State ablations show extra compute/memory is not enough; gradient analysis routes future-loss onto the input stream. Code https://github.com/lil-lab/sps.

Keywords

sps · two-stream · kv-cache · state-prediction · fineweb-edu · cornell · harvard

Topics

transformer architecture, language modeling

Research notes

  • Primary: arxiv abs (default nonexclusive-distrib 1.0, cs.LG/CL). Cornell LIL Lab; Brantley at Harvard. Correspondence giovanni@cs.cornell.edu. Code https://github.com/lil-lab/sps (23 stars at check); license not stated on the GitHub landing page, so left blank. Discord posted AlphaXiv 2607.01218; canonical abs recorded. HF paper page 12 upvotes, org lil-lab; no linked models/datasets. Trains on public FineWeb-Edu; no new corpus, so no datasets_local row.