Spectral Scaling Laws of Muon
- Type
- other
- Venue
- arXiv / MIT
Summary
Tracks Frobenius-normalized momentum singular-value quantiles across GPT-2-style models 77M–2.8B trained with Muon on FineWeb. After a short burn-in, quantiles stabilize by layer type and follow power laws in model size M with exponents from about M^{-0.25} (mid-depth) to M^{-0.96} (final MLP). Rank-p ablations show orthonormalizing the top ~50% of directions nearly matches full Muon, while top 10% is ~50% less token-efficient. Extrapolating to 300B, standard 5-step NanoGPT NS still covers mid-late Q, but final O falls into the failure regime and wants a 10-step DeepSeek-V4-style map. No official code repo on the abs page.
Keywords
muon · newton-schulz · spectral-scaling · orthonormalization · mit · fineweb
Topics
optimization, Muon, Newton-Schulz, scaling laws
Research notes
- Primary: arxiv abs. Correspondence gagmag@mit.edu. MIT. Discord posted tweet + pdf; canonical abs recorded. HF paper page exists. No official training-code repo on abs (experiments use modded-nanogpt). ArXiv license widget not visible in converted abs HTML, so license left blank. Uses public FineWeb; no new corpus, so no datasets_local row.