Emergent One-Third Scaling Law as Attention Tries to Concentrate
- Type
- paper
- Venue
- arXiv:2609.32100 (cs.LG), submitted 26 Sep 2026
- Year
- 2026
- Source
- arxiv
- Access
- free
- Language
- English
- Added
- 2026-10-02
- Verified
- 2026-10-02
Summary
Proposes an origin for neural scaling laws: any softmax learning peaked distributions, regardless of its position in the model, develops logit magnitudes that grow as a power law with exponent 1/3, becoming a training bottleneck whose loss contribution decays with the same 1/3 exponent — so total loss obeys 1/3 scaling whenever at least one softmax learns peaked distributions. Toy models show the effect for both output selection and attention-like gating, with a 'loss universality' result (any reasonable loss yields the same law; nonlinearity is the key ingredient). In real LLMs (Pythia, OLMo), loss falls and attention's internal logits grow with exponents near 1/3 while the output head shows no such growth, pointing to attention heads — not the LM head — as the bottleneck driving 1/3 loss scaling. Bypassing the output softmax with MSE training, and using sigmoid/tanh gates, preserves the law.
Keywords
one-third scaling law · neural scaling laws · attention · softmax · peaked distributions · logit growth · loss universality · Pythia · OLMo · massive activations
Topics
neural scaling laws, attention, softmax, power laws, training dynamics, massive activations
Research notes
- Discovery: @YizhouLiu0 X thread (12 parts) 2026-10-01 (https://x.com/YizhouLiu0/status/2105675423157748058)
- Code: https://github.com/liuyz0/AttnScaling
- Builds on the authors' prior work 'Universal One-third Time Scaling in Learning Peaked Distributions' (arXiv:2602.03685), which showed focusing on a particular output produces the 1/3 law; the new puzzle was why focusing attention inside the model gives the same law despite layers between attention and prediction
- Toy model with two jobs (choosing what to focus on, choosing what to output): when either part learns to focus sharply, loss falls as a 1/3 power law; training only the gate reproduces the law in both directions (prediction error falls, internal focus scores grow, both ~1/3)
- 'Loss universality' proof: any reasonable loss leads to the same 1/3 law, nonlinearity is the key; layers after attention effectively create a more complicated loss over attention probabilities, so middle layers produce 1/3 end-loss scaling
- Real LLMs: Pythia and OLMo both show loss falling and attention internal logits growing with exponents near 1/3, output head shows no growth; bypassing output softmax (MSE training) keeps the law while the output head settles and attention scores keep growing
- Beyond softmax: sigmoid and tanh gates show the same 1/3 law — nonlinear saturation may be the common ingredient; thread speculates GLU/MoE-style architectures could lower achievable error yet add more such bottlenecks
- Affiliations per announcement: Sara Kangaslahti (Harvard), Jeff Gore (MIT)
- License: CC BY 4.0