← Back to explorer

Spectral Scaling Laws of Muon

Type
other
Venue
arXiv / MIT

Summary

Tracks Frobenius-normalized momentum singular-value quantiles across GPT-2-style models 77M–2.8B trained with Muon on FineWeb. After a short burn-in, quantiles stabilize by layer type and follow power laws in model size M with exponents from about M^{-0.25} (mid-depth) to M^{-0.96} (final MLP). Rank-p ablations show orthonormalizing the top ~50% of directions nearly matches full Muon, while top 10% is ~50% less token-efficient. Extrapolating to 300B, standard 5-step NanoGPT NS still covers mid-late Q, but final O falls into the failure regime and wants a 10-step DeepSeek-V4-style map. No official code repo on the abs page.

Keywords

muon · newton-schulz · spectral-scaling · orthonormalization · mit · fineweb

Topics

optimization, Muon, Newton-Schulz, scaling laws

Research notes

  • Primary: arxiv abs. Correspondence gagmag@mit.edu. MIT. Discord posted tweet + pdf; canonical abs recorded. HF paper page exists. No official training-code repo on abs (experiments use modded-nanogpt). ArXiv license widget not visible in converted abs HTML, so license left blank. Uses public FineWeb; no new corpus, so no datasets_local row.