← Back to explorer

Decoding Looped Transformers Better for (Almost) Free

Type
paper
Venue
arXiv:2610.02185 (cs.LG), submitted 1 Oct 2026
Year
2026
Source
arxiv
Access
free
Language
English
Added
2026-10-02
Verified
2026-10-02

Summary

Looped Transformers repeatedly execute a shared block across recurrent loops; each loop yields an intermediate representation decodable for the same next token, but standard decoding discards earlier states. Because earlier loops embody less computation, recurrence inherently supplies aligned weak-and-strong prediction pairs without auxiliary models or external training. LoopCD is a training-free contrastive decoding framework that guides token selection by contrasting the final prediction with an earlier recurrent pass — in logit space with one extra output pass (LoopCD-Logits) or in hidden-state space with zero output overhead (LoopCD-Hidden). Across four looped Transformer families it delivers consistent gains at full recurrent depth: LoopCD-Logits raises Ouro-2.6B-Thinking's AIME 2024 pass@1 from 61.88% to 73.33%, and LoopCD-Hidden lifts Huginn's HumanEval pass@1 from 22.56% to 31.71%. The gains enable halving the number of recurrent loops while still matching or exceeding full-depth unguided baselines, reducing forward FLOPs by 22.5%-48.2%.

Keywords

LoopCD · looped Transformers · contrastive decoding · Ouro · Huginn · recurrent loops · inference FLOPs · training-free

Topics

looped Transformers, contrastive decoding, inference efficiency, recurrence, Ouro, Huginn

Research notes

  • Discovery: @arankomatsuzaki X post 2026-10-02 (https://x.com/arankomatsuzaki/status/2105888510482034876)
  • Apple paper (first-page screenshot; Weihao Liu correspondence wliu681@uic.edu, work done at Apple internship; Ruixiang Zhang ruixiangz@apple.com)
  • Method: contrast final prediction against an earlier recurrent pass — LoopCD-Logits (logit space, one extra output pass) or LoopCD-Hidden (hidden-state space, zero output overhead); no auxiliary models, no training
  • Results across four looped Transformer families: Ouro-2.6B-Thinking AIME 2024 pass@1 61.88% -> 73.33% (LoopCD-Logits); Huginn HumanEval pass@1 22.56% -> 31.71% (LoopCD-Hidden); halving recurrent loops still matches or exceeds full-depth unguided baselines, cutting forward FLOPs 22.5%-48.2%
  • Related catalog entry: Papers 928 (Scaling Laws for Looped Mixture of Experts) — looped-architecture companion
  • License not stated on the arXiv page (generic 'view license' shown)