← Back to explorer

High-Dimensional Learning Dynamics of Attention-Indexed Models

Type
paper
Venue
arXiv (2026-09-03), cs.LG / stat.ML; arXiv non-exclusive distribution license
Year
2026
Source
paper
Access
public
Language
en
Added
2026-09-30
Verified
2026-09-30

Summary

ML theory paper on "attention-indexed models," a framework for multi-layer and multi-head attention architectures in the high-dimensional limit. Core result: the population-loss landscape collapses into a finite set of trace order parameters, yet under online SGD the parameter trajectories are governed by an infinite hierarchy of coupled matrix moments (each SGD step mixes S_k with the gradient tensor, so degree-w moments couple to degree-(w+1); degree-M truncations approximate the exact trajectory with exponential precision O((C/s0)^M)). Parameterization is an active regularizer: direct optimization of an attention matrix S in R^(d x d) can stay trapped on an uninformative manifold; tied attention (S = WW^T) forces a positive initial trace and yields automatic symmetry-breaking with weak recovery in Theta(d^2 log d) samples; untied attention (S = UV^T) shows a fast-slow relaxation (fast pre-activation mean drift toward a critical manifold, then slow feature recovery only if the fast boundary layer broke student symmetry). Caveats: assumes online SGD with one-pass streaming Gaussian data (no finite-sample ERM or multi-pass token reuse) and analyzes effective bilinear interactions rather than full end-to-end multi-layer QKV co-evolution.

Keywords

attention · optimization theory · learning dynamics · ML theory

Topics

attention, optimization theory, learning dynamics, ML theory

Research notes

  • Discovery: Grigory Sapunov (@che_shr_cat, verified, arXiviq ML-theory newsletter) 2026-09-30 10-tweet explainer thread: https://x.com/che_shr_cat/status/2105377511593103657?s=20
  • Full breakdown: https://arxiviq.substack.com/p/high-dimensional-learning-dynamics
  • License: arXiv non-exclusive distribution license (no open CC license on the arXiv page).
  • No code or dataset links given.
  • Connects to the collection's attention-optimization and pretraining-stability theory entries.