High-Dimensional Learning Dynamics of Attention-Indexed Models
- Type
- paper
- Venue
- arXiv (2026-09-03), cs.LG / stat.ML; arXiv non-exclusive distribution license
- Year
- 2026
- Source
- paper
- Access
- public
- Language
- en
- Added
- 2026-09-30
- Verified
- 2026-09-30
Summary
ML theory paper on "attention-indexed models," a framework for multi-layer and multi-head attention architectures in the high-dimensional limit. Core result: the population-loss landscape collapses into a finite set of trace order parameters, yet under online SGD the parameter trajectories are governed by an infinite hierarchy of coupled matrix moments (each SGD step mixes S_k with the gradient tensor, so degree-w moments couple to degree-(w+1); degree-M truncations approximate the exact trajectory with exponential precision O((C/s0)^M)). Parameterization is an active regularizer: direct optimization of an attention matrix S in R^(d x d) can stay trapped on an uninformative manifold; tied attention (S = WW^T) forces a positive initial trace and yields automatic symmetry-breaking with weak recovery in Theta(d^2 log d) samples; untied attention (S = UV^T) shows a fast-slow relaxation (fast pre-activation mean drift toward a critical manifold, then slow feature recovery only if the fast boundary layer broke student symmetry). Caveats: assumes online SGD with one-pass streaming Gaussian data (no finite-sample ERM or multi-pass token reuse) and analyzes effective bilinear interactions rather than full end-to-end multi-layer QKV co-evolution.
Keywords
attention · optimization theory · learning dynamics · ML theory
Topics
attention, optimization theory, learning dynamics, ML theory
Research notes
- Discovery: Grigory Sapunov (@che_shr_cat, verified, arXiviq ML-theory newsletter) 2026-09-30 10-tweet explainer thread: https://x.com/che_shr_cat/status/2105377511593103657?s=20
- Full breakdown: https://arxiviq.substack.com/p/high-dimensional-learning-dynamics
- License: arXiv non-exclusive distribution license (no open CC license on the arXiv page).
- No code or dataset links given.
- Connects to the collection's attention-optimization and pretraining-stability theory entries.