← Back to explorer

Transformers Stop Thinking Too Early, and a Tiny LoRA Fixes It

Type
paper
Venue
arXiv (2026-09-29), cs.AI; code + interactive demo available
Year
2026
Source
paper
Access
public
Language
en
Added
2026-09-30
Verified
2026-09-30

Summary

Pretrained transformers use little of their depth to follow references in context: 13 base models reliably follow only 1.4-3.6 lines, and extra pretrained loops add little. A task-trained rank-8 LoRA at one early layer extends computation with all model weights frozen: Qwen3-8B improves from 15.5% to 99% exact accuracy on 24-line chains; a longer-trained LoRA reaches 50 lines; Ouro-1.4B reaches 60 lines after four loops and at least 160 after eight. Mechanism: the LoRA starts a relay -- program lines pass on their chain identity through a short range of middle layers, frozen heads read progressively further up the chain, and removing parent-line attention stops the relay. A frozen-model measurement locates the last useful intervention layer within tolerance in 3 of 4 held-out models. Task-specific LoRAs also improve MuSiQue. Takeaway: default answers understate the computation accessible through a tiny edit. Code and an interactive demo at https://lunamos.github.io/stop-thinking-too-early/

Keywords

transformers · LoRA · reasoning depth · multi-hop retrieval

Topics

transformers, LoRA, reasoning depth, multi-hop retrieval

Research notes

  • Discovered via arXivBangers (arxivb.org) "Certified banger", awarded 2026-09-30, 88/100 editorial score: https://arxivb.org/abs/2609.36585
  • Code + interactive demo: https://lunamos.github.io/stop-thinking-too-early/
  • Connects to the collection's long-context, LoRA, and multi-hop reasoning entries.