← Back to explorer

A Theoretical Analysis of Generalization Dynamics in Neural Networks under Gradient Descent with Weight Decay

Type
paper
Venue
arXiv (cs.LG), submitted 2026-09-07
Year
2026
Source
arxiv
Access
public
Language
en
Added
2026-09-30
Verified
2026-09-30

Summary

Theoretical framework characterizing how data, architecture, and training dynamics jointly shape generalization throughout training. Studies neural networks trained with L2 loss by gradient descent with weight decay; proves GD converges to a neighbourhood of the global minimizers of the empirical loss, then partitions the input space into data-centered cells and decomposes population risk into three explicit terms: E_opt (optimization error on training data), E_PV (propagation-variation error: how predictions oscillate within each cell), and E_A (approximation error: how far ground truth is from the function class). The culprit behind memorization without generalization is E_PV: E_opt vanishes early while E_PV stays large. Under approximate Euler homogeneity, layerwise variation satisfies a differential inequality with explicit contraction rate lambda_l >= 1 - 2(1 + L) eta_2 lambda-bar, bounding E_PV by the initial hypothesis variation decaying exponentially in steps at a rate proportional to weight decay. Yields a closed-form grokking residency time (Steps ~ Omega(log(Initial_Var / budget) / (lambda-bar * eta_2 * sum L_l))): weight decay accelerates generalization, deeper networks contract faster, smoother initialization shortens the delay. Bounds are qualitative -- constants depend on suprema over parameter balls and loss smoothness, so they give rates and scaling laws, not exact step counts.

Keywords

generalization theory · grokking · weight decay · optimization · memorization

Topics

generalization theory, grokking, weight decay, optimization

Research notes

  • Discovery: Grigory Sapunov (@che_shr_cat, verified) 2026-09-30 thread: https://x.com/che_shr_cat/status/2105179516738175168
  • Thread takeaways: grokking is not magic and flat minima do not explain it -- weight decay is the active engine contracting the hypothesis between data points, driving memorization-to-generalization transition; delayed generalization is the signature of a function being smoothed by its own training dynamics.
  • Sapunov's full technical breakdown: https://arxiviq.substack.com/p/a-theoretica...
  • Connects to the collection's generalization-theory, grokking, and optimization entries.