← Back to explorer

Emergent One-Third Scaling Law as Attention Tries to Concentrate

Type
paper
Venue
arXiv:2609.32100 (cs.LG), submitted 26 Sep 2026
Year
2026
Source
arxiv
Access
free
Language
English
Added
2026-10-02
Verified
2026-10-02

Summary

Proposes an origin for neural scaling laws: any softmax learning peaked distributions, regardless of its position in the model, develops logit magnitudes that grow as a power law with exponent 1/3, becoming a training bottleneck whose loss contribution decays with the same 1/3 exponent — so total loss obeys 1/3 scaling whenever at least one softmax learns peaked distributions. Toy models show the effect for both output selection and attention-like gating, with a 'loss universality' result (any reasonable loss yields the same law; nonlinearity is the key ingredient). In real LLMs (Pythia, OLMo), loss falls and attention's internal logits grow with exponents near 1/3 while the output head shows no such growth, pointing to attention heads — not the LM head — as the bottleneck driving 1/3 loss scaling. Bypassing the output softmax with MSE training, and using sigmoid/tanh gates, preserves the law.

Keywords

one-third scaling law · neural scaling laws · attention · softmax · peaked distributions · logit growth · loss universality · Pythia · OLMo · massive activations

Topics

neural scaling laws, attention, softmax, power laws, training dynamics, massive activations

Research notes

  • Discovery: @YizhouLiu0 X thread (12 parts) 2026-10-01 (https://x.com/YizhouLiu0/status/2105675423157748058)
  • Code: https://github.com/liuyz0/AttnScaling
  • Builds on the authors' prior work 'Universal One-third Time Scaling in Learning Peaked Distributions' (arXiv:2602.03685), which showed focusing on a particular output produces the 1/3 law; the new puzzle was why focusing attention inside the model gives the same law despite layers between attention and prediction
  • Toy model with two jobs (choosing what to focus on, choosing what to output): when either part learns to focus sharply, loss falls as a 1/3 power law; training only the gate reproduces the law in both directions (prediction error falls, internal focus scores grow, both ~1/3)
  • 'Loss universality' proof: any reasonable loss leads to the same 1/3 law, nonlinearity is the key; layers after attention effectively create a more complicated loss over attention probabilities, so middle layers produce 1/3 end-loss scaling
  • Real LLMs: Pythia and OLMo both show loss falling and attention internal logits growing with exponents near 1/3, output head shows no growth; bypassing output softmax (MSE training) keeps the law while the output head settles and attention scores keep growing
  • Beyond softmax: sigmoid and tanh gates show the same 1/3 law — nonlinear saturation may be the common ingredient; thread speculates GLU/MoE-style architectures could lower achievable error yet add more such bottlenecks
  • Affiliations per announcement: Sara Kangaslahti (Harvard), Jeff Gore (MIT)
  • License: CC BY 4.0