← Back to explorer

Pretraining Latent Information Feedback Transformers with Teacher Supervision

Type
paper
Venue
arXiv:2609.38149 (cs.CL), submitted 29 Sep 2026
Year
2026
Source
arxiv
Access
free
Language
English
Added
2026-10-01
Verified
2026-10-01

Summary

Introduces LIFT (Latent Information Feedback Transformer), which removes the decoded-token bottleneck that normally prevents deep-layer representations from feeding back to shallower layers during generation. Each input token is paired with an information-dense teacher state derived from an off-the-shelf pretrained LM's next-token distribution, and the model learns to predict both the next token and the next state; because states are precomputed, pretraining stays fully parallel while inference feeds back the model's own predicted states. Across 135M-1B models, LIFT beats standard Transformers on language modeling, downstream reasoning, and procedural tasks under token-matched budgets, and a tiny LIFT beats same-size Transformers trained on 8x more data on a state-tracking task.

Keywords

LIFT · Latent Information Feedback Transformer · teacher supervision · deep-to-shallow feedback · latent state propagation · parallel pretraining · state tracking

Topics

Transformer architecture, latent state feedback, teacher supervision, pretraining, state tracking

Research notes

  • Discovery: @megamor2 X post 2026-10-01 (https://x.com/megamor2/status/2105720584655237197)
  • Teacher states are derived from the next-token distribution of an off-the-shelf pretrained LM; model is trained to predict both next token and next state, turning recurrent-state learning into a teacher-forced prediction problem; at inference the model's own predicted states are fed back with minor overhead that decreases with model size
  • Experiments at 135M-1B parameters: outperforms standard Transformers and baselines on language modeling, downstream reasoning, and procedural tasks under token-matched budget, on par with or ahead of compute-matched Transformers; controlled state-tracking study: tiny LIFT outperforms same-size Transformers trained on 8x more data, even when trained with states from a Transformer that fails the task
  • Affiliations (paper first page): Dor Tirosh and Mor Geva, Blavatnik School of Computer Science and AI, Tel Aviv University; Ido Amos, The Hebrew University of Jerusalem
  • Announcement says detailed post + visualizations + code + models coming soon; no code/model release yet
  • License: CC BY 4.0