Transformers Stop Thinking Too Early, and a Tiny LoRA Fixes It
- Type
- paper
- Venue
- arXiv (2026-09-29), cs.AI; code + interactive demo available
- Year
- 2026
- Source
- paper
- Access
- public
- Language
- en
- Added
- 2026-09-30
- Verified
- 2026-09-30
Summary
Pretrained transformers use little of their depth to follow references in context: 13 base models reliably follow only 1.4-3.6 lines, and extra pretrained loops add little. A task-trained rank-8 LoRA at one early layer extends computation with all model weights frozen: Qwen3-8B improves from 15.5% to 99% exact accuracy on 24-line chains; a longer-trained LoRA reaches 50 lines; Ouro-1.4B reaches 60 lines after four loops and at least 160 after eight. Mechanism: the LoRA starts a relay -- program lines pass on their chain identity through a short range of middle layers, frozen heads read progressively further up the chain, and removing parent-line attention stops the relay. A frozen-model measurement locates the last useful intervention layer within tolerance in 3 of 4 held-out models. Task-specific LoRAs also improve MuSiQue. Takeaway: default answers understate the computation accessible through a tiny edit. Code and an interactive demo at https://lunamos.github.io/stop-thinking-too-early/
Keywords
transformers · LoRA · reasoning depth · multi-hop retrieval
Topics
transformers, LoRA, reasoning depth, multi-hop retrieval
Research notes
- Discovered via arXivBangers (arxivb.org) "Certified banger", awarded 2026-09-30, 88/100 editorial score: https://arxivb.org/abs/2609.36585
- Code + interactive demo: https://lunamos.github.io/stop-thinking-too-early/
- Connects to the collection's long-context, LoRA, and multi-hop reasoning entries.