← Back to explorer

Invariant Learning Dynamics of Transformers in Inductive Reasoning Tasks

Type
paper
Venue
arXiv / ETH Zurich
Year
2026
Source
arxiv
Access
free
Language
en
Added
2026-08-14T16:54:04Z
Verified
2026-08-14T16:54:04Z

Summary

Defines the Invariant Manifold of Inductive Reasoning (IMIR): a low-dimensional subspace that gradient descent never leaves on a block-list task class unifying in-context n-grams, k-hop induction, and associative recall. QK and output weights stay in a span of interpretable token/position selection and action bases. In a 2-layer 1-head model the IMIR is 28-dimensional; ICL (alpha, beta, gamma induction head) competes with IWL (delta). IWL emerges first and multiplies ICL gradients by a data-dependent factor; burstiness multiplies the ICL gradient. In 3-layer models, initialization selects among three induction-head circuits with sharp cooperative/competitive phase boundaries. Projecting trained weights onto the IMIR recovers two-hop induction circuits including extra aiding directions. Analysis is attention-only with fixed orthogonal OV maps and merged QK; FFN/LN extensions are sketched. No official code on the abs page.

Keywords

imir · invariant-manifold · induction-heads · icl · iwl · circuit-competition · eth-zurich

Topics

interpretability, learning dynamics, transformers

Research notes

  • Primary: arxiv abs (CC BY 4.0, cs.LG). Correspondence tiberiu@musat.ai, tiago.pimentel@inf.ethz.ch. ETH Zurich; Zucchet at Stanford. No official code on abs. HF has no paper page (404). Discord posted abs link. Synthetic tasks only; no datasets_local row.