Pretraining Latent Information Feedback Transformers with Teacher Supervision
- Type
- paper
- Venue
- arXiv:2609.38149 (cs.CL), submitted 29 Sep 2026
- Year
- 2026
- Source
- arxiv
- Access
- free
- Language
- English
- Added
- 2026-10-01
- Verified
- 2026-10-01
Summary
Introduces LIFT (Latent Information Feedback Transformer), which removes the decoded-token bottleneck that normally prevents deep-layer representations from feeding back to shallower layers during generation. Each input token is paired with an information-dense teacher state derived from an off-the-shelf pretrained LM's next-token distribution, and the model learns to predict both the next token and the next state; because states are precomputed, pretraining stays fully parallel while inference feeds back the model's own predicted states. Across 135M-1B models, LIFT beats standard Transformers on language modeling, downstream reasoning, and procedural tasks under token-matched budgets, and a tiny LIFT beats same-size Transformers trained on 8x more data on a state-tracking task.
Keywords
LIFT · Latent Information Feedback Transformer · teacher supervision · deep-to-shallow feedback · latent state propagation · parallel pretraining · state tracking
Topics
Transformer architecture, latent state feedback, teacher supervision, pretraining, state tracking
Research notes
- Discovery: @megamor2 X post 2026-10-01 (https://x.com/megamor2/status/2105720584655237197)
- Teacher states are derived from the next-token distribution of an off-the-shelf pretrained LM; model is trained to predict both next token and next state, turning recurrent-state learning into a teacher-forced prediction problem; at inference the model's own predicted states are fed back with minor overhead that decreases with model size
- Experiments at 135M-1B parameters: outperforms standard Transformers and baselines on language modeling, downstream reasoning, and procedural tasks under token-matched budget, on par with or ahead of compute-matched Transformers; controlled state-tracking study: tiny LIFT outperforms same-size Transformers trained on 8x more data, even when trained with states from a Transformer that fails the task
- Affiliations (paper first page): Dor Tirosh and Mor Geva, Blavatnik School of Computer Science and AI, Tel Aviv University; Ido Amos, The Hebrew University of Jerusalem
- Announcement says detailed post + visualizations + code + models coming soon; no code/model release yet
- License: CC BY 4.0