Invariant Learning Dynamics of Transformers in Inductive Reasoning Tasks
- Type
- paper
- Venue
- arXiv / ETH Zurich
- Year
- 2026
- Source
- arxiv
- Access
- free
- Language
- en
- Added
- 2026-08-14T16:54:04Z
- Verified
- 2026-08-14T16:54:04Z
Summary
Defines the Invariant Manifold of Inductive Reasoning (IMIR): a low-dimensional subspace that gradient descent never leaves on a block-list task class unifying in-context n-grams, k-hop induction, and associative recall. QK and output weights stay in a span of interpretable token/position selection and action bases. In a 2-layer 1-head model the IMIR is 28-dimensional; ICL (alpha, beta, gamma induction head) competes with IWL (delta). IWL emerges first and multiplies ICL gradients by a data-dependent factor; burstiness multiplies the ICL gradient. In 3-layer models, initialization selects among three induction-head circuits with sharp cooperative/competitive phase boundaries. Projecting trained weights onto the IMIR recovers two-hop induction circuits including extra aiding directions. Analysis is attention-only with fixed orthogonal OV maps and merged QK; FFN/LN extensions are sketched. No official code on the abs page.
Keywords
imir · invariant-manifold · induction-heads · icl · iwl · circuit-competition · eth-zurich
Topics
interpretability, learning dynamics, transformers
Research notes
- Primary: arxiv abs (CC BY 4.0, cs.LG). Correspondence tiberiu@musat.ai, tiago.pimentel@inf.ethz.ch. ETH Zurich; Zucchet at Stanford. No official code on abs. HF has no paper page (404). Discord posted abs link. Synthetic tasks only; no datasets_local row.