Demystifying Manifold Constraints in LLM Pre-training
- Type
- other
- Venue
- arXiv / Rice University / Columbia University
Summary
Introduces MACRO, a single-loop Riemannian spectral-SGD (Muon-style msign on the tangent-space gradient plus retraction) with O(T^{-1/4}) stochastic nonconvex rate. Spectral sphere bounds worst-case activation scale; Frobenius sphere bounds average-case (radius r√D_out). Without learnable RMSNorm, Muon NaNs at large LR while MACRO-spec stays stable (330M Qwen3-like val 2.739 vs normalized MACRO-spec 2.714). Constraints lock relative LR η_rel=cη and rotational equilibrium from step 1, replacing decoupled weight decay. On standard (RMSNorm) Qwen3-like 120M–1B above Chinchilla, MACRO-spec matches or slightly beats Muon/MuonH/SSO with exact Riemannian updates. No official code on abs.
Keywords
macro · muon · manifold-constraints · spectral-sphere · frobenius-sphere · rmsnorm · weight-decay · rice · columbia
Topics
optimizers, manifold constraints, LLM pretraining
Research notes
- Primary: arxiv abs (cs.LG; also cs.AI, math.OC). ArXiv HTML states perpetual non-exclusive license. An/Li equal contrib. An/Ma Rice (kang.an/shiqian.ma@rice.edu); Li Independent (jasonljx96@gmail.com); Goldfarb Columbia. No official code on abs. HF has no paper page (API 404). Discord posted abs. Trains Qwen3-like models on standard pretraining data; no new corpus, so no datasets_local row. License field left blank per catalog convention.