← Back to explorer

The Loss Does Not See the Basis, but Adam Does

Type
paper
Venue
arXiv
Year
2026
Source
arxiv
Access
free
Language
en
Added
2026-08-14T16:19:42Z
Verified
2026-08-14T16:19:42Z

Summary

Shows gradient descent on W=UV^T is biased toward low-rank interpolants while Adam is not, tracing the gap to loss gauge symmetry (U,V)->(UQ,VQ). Equivariant methods (GD, momentum, shared-scalar Adam, Muon, Shampoo) can inherit gradient-flow low-rank bias; coordinate-wise Adam/RMSProp/Lion/Adafactor cannot. Sorts nine update rules on underdetermined matrix sensing; a p-dial from coordinate-wise to shared-scalar Adam restores the bias monotonically. In transformers, Adam splits two gauge-equivalent inits at step 1 (per-head W_Q^T W_K 56% apart). On two hyperspectral datasets at matched train loss, GD cuts held-out error 43-44% vs Adam at lowest sampling density.

Keywords

adam · muon · shampoo · gauge-equivariance · implicit-bias · factored-models · matrix-sensing · attention

Topics

optimization, implicit bias, transformers

Research notes

  • Primary: arxiv abs. Code https://github.com/idevender/loss-basis-adam project https://idevender.github.io/loss-basis-adam/. HF paper page has 7 upvotes. Did not copy CC-BY; arXiv default license.