One LR Doesn't Fit All: Heavy-Tail Guided Layerwise Learning Rates for LLMs
- Type
- other
- Venue
- arXiv / SIAT CAS / Peng Cheng Laboratory / University of Surrey / MPI-IS / ELLIS Institute Tübingen
Summary
Uniform LR ignores Transformer heterogeneity. LLR fits a power-law to each layer's weight-correlation ESD (PL_Alpha_Hill) and maps weaker heavy tails (embeddings/FFN) to larger LRs and stronger tails (attention) to smaller ones, with embedding pinned at the upper bound, a soft LR switch, and updates only in the first 20% of tokens. Transfers the uniform baseline's near-optimal global LR. LLaMa-1B FineWeb: zero-shot avg 47.09→49.02; 3B 48.58→50.61; up to 1.5x token speedup. Beats LARS/LAMB/Sharpness/TempBalance/AlphaDecay; also helps Muon and GPT-nano. Code https://github.com/hed-ucas/Layer-wise-Learning-Rate.
Keywords
layerwise-lr · ht-sr · llama · muon · fineweb · icml · mpi-is · pl-alpha-hill
Topics
LLM pretraining, learning rates, HT-SR, optimization
Research notes
- Primary: arxiv abs (cs.LG/AI). SIAT CAS / Peng Cheng Lab / UCAS / Tübingen / Surrey / MPI-IS / ELLIS Institute Tübingen. Correspondence l.yin@surrey.ac.uk, sliu@tue.ellis.eu. Code https://github.com/hed-ucas/Layer-wise-Learning-Rate (9 stars at check; no LICENSE file). HF paper page 1 upvote. Discord posted PDF. Pretrains on public FineWeb; no new corpus, so no datasets_local row. ArXiv license widget not visible in converted abs HTML, so paper license left blank.