The Loss Does Not See the Basis, but Adam Does
- Type
- paper
- Venue
- arXiv
- Year
- 2026
- Source
- arxiv
- Access
- free
- Language
- en
- Added
- 2026-08-14T16:19:42Z
- Verified
- 2026-08-14T16:19:42Z
Summary
Shows gradient descent on W=UV^T is biased toward low-rank interpolants while Adam is not, tracing the gap to loss gauge symmetry (U,V)->(UQ,VQ). Equivariant methods (GD, momentum, shared-scalar Adam, Muon, Shampoo) can inherit gradient-flow low-rank bias; coordinate-wise Adam/RMSProp/Lion/Adafactor cannot. Sorts nine update rules on underdetermined matrix sensing; a p-dial from coordinate-wise to shared-scalar Adam restores the bias monotonically. In transformers, Adam splits two gauge-equivalent inits at step 1 (per-head W_Q^T W_K 56% apart). On two hyperspectral datasets at matched train loss, GD cuts held-out error 43-44% vs Adam at lowest sampling density.
Keywords
adam · muon · shampoo · gauge-equivariance · implicit-bias · factored-models · matrix-sensing · attention
Topics
optimization, implicit bias, transformers
Research notes
- Primary: arxiv abs. Code https://github.com/idevender/loss-basis-adam project https://idevender.github.io/loss-basis-adam/. HF paper page has 7 upvotes. Did not copy CC-BY; arXiv default license.