← Back to explorer

Towards Looped Models Done Right — Part II: Rethinking at Fixed Points

Type
paper
Venue
alphaXiv 2610.looped-models-fixed-points, 2 Oct 2026 (arXiv release pending)
Year
2026
Source
web
Access
free
Language
English
Added
2026-10-02
Verified
2026-10-02

Summary

Looped LMs need not pay for every loop: they can scale by adding FLOPs at constant memory, and the key is the loop's fixed point. Three fixed-point ideas: (1) Fixed-Point Shortcuts — what fixed points buy in training (pre-train, post-train) and inference (prefill, decoding). Terminal KV sharing: each prefix token keeps only its last-loop KV (reuse the KV from the last loop), discarding intermediate loops; with depth sampling it keeps 4 of 12 KV banks with at most 0.85% perplexity increase from 100M (21.5B tokens) to 1.6B (344B tokens). A model trained at one fixed depth breaks under sharing (red rows) — sharing works only when states settle (95% of tokens stable by loop 8) because deep draws train on a nearly settled prefix KV. Proven: both training-read and sharing-read reach the same fixed point and same next-token prediction when the prefix KV converges and each token's update is a contraction (Lemma 1). (2) Learned depth prior — prediction feedback moves draws toward depths that predict well, an entropy term keeps the deep draws that form fixed points, a budget term holds mean depth; at 1.6B with the shared cache it averages 60.0 (level with the 59.9 fixed-depth reaches with a 3x larger cache), while at 1.6B/344B tokens the looped model beats a 4-block Transformer by 8.9 average points and lands within 1.5 pts of a 12-block model needing 3x memory. (3) OrthoInj (orthogonal input injection) — projects out the injected-input component so input conditioning stays constant across loops; lowest loss and val PPL, highest downstream average at every scale. Also: terminal KV sharing lets a distilled student prefill by predicting the endpoint (1.79x faster prefill of an 8K prompt at 1.6B, keeps 93% of teacher downstream score); truncated BPTT with depth sampling matches full backprop (backprop window b=4 under the PLN-5 prior) and covers both pre-fixed-point training and the endpoint gradient; fixed-point reuse for RL skips the redundant second pass by reusing the rollout's final states and applying the block once (2x faster scoring/backward on GSM8K and MBPP+, pass@1 within run-to-run variation). Limits: models reach 1.6B, one training seed per configuration, larger scales and real-world feasibility untested.

Keywords

looped LMs · fixed points · KV cache sharing · truncated BPTT · learned depth prior · orthogonal injection · weight tying

Topics

looped LMs, weight tying, fixed points, KV cache compression, truncated BPTT, depth prior, orthogonal input injection, test-time compute

Research notes

  • Discovery: @huskydogewoof (Benhao Huang, CMU/IFM-AI) long X thread 2026-10-02 5:14 PM ET (https://x.com/huskydogewoof/status/2106130737007091939)
  • Paper on alphaXiv ('arXiv to be released'); follow-up to Part I, already cataloged as paper 670 ('Topology, Input Injection, Recurrent-State Design')
  • Code/data/checkpoints fully open: github.com/ifm-ai/xllm-loop (official code for Parts I and II; 49 HF checkpoints referenced in README)
  • Teammates per post: @Chufan_Shi, @Chen94751623484 (Junlin Chen), @Shicheng_Wen, @MaxMa1987, @waterluffy, @ericxing, @IFM_AI
  • Terminal KV sharing idea also explored in Huginn (@jonasgeiping); this paper adds the why (contraction proof, Lemma 1)