← Back to explorer

Training-Free Looped Transformers

Type
other
Venue
arXiv / University of Texas at Austin

Summary

Lightweight wrapper re-applies a contiguous mid-stack block of a frozen LM. Naive block reapplication usually degrades; treating a pre-norm block as a forward Euler step and replacing one large update with damped/RK sub-steps helps. Single out-of-the-box recipe (3-stage Runge–Kutta at the mid 4 layers; block-mode for dense, layer-mode for MoE so expert routing does not thrash) over 7 families and 45 (model, benchmark) cells. Qwen3-4B-Instruct +2.64 pp MMLU-Pro and +2.01 GPQA-Main; Qwen1.5-MoE-A2.7B-Chat +2.30 ARC-Challenge; Qwen3-30B-A3B-Instruct +1.14 CommonsenseQA; Moonlight-16B-A3B-Instruct +1.20 OpenBookQA. ~20k H100 hours. No official code on abs.

Keywords

looped-transformer · training-free · runge-kutta · moe · qwen3 · moonlight · ut-austin · inference-time-compute

Topics

looped transformers, inference-time compute, ODE / numerical integration

Research notes

  • Primary: arxiv abs (cs.LG / math.NA / stat.ML). UT Austin; Chen/Li equal contrib. No official code on abs. HF paper page 2 upvotes; no linked models/datasets. Discord posted AlphaXiv 2605.23872. Uses public checkpoints and MC benches; no new corpus, so no datasets_local row. ArXiv HTML states perpetual non-exclusive license; license field left blank per catalog convention when the abs widget is not used.