A Theoretical Analysis of Generalization Dynamics in Neural Networks under Gradient Descent with Weight Decay
- Type
- paper
- Venue
- arXiv (cs.LG), submitted 2026-09-07
- Year
- 2026
- Source
- arxiv
- Access
- public
- Language
- en
- Added
- 2026-09-30
- Verified
- 2026-09-30
Summary
Theoretical framework characterizing how data, architecture, and training dynamics jointly shape generalization throughout training. Studies neural networks trained with L2 loss by gradient descent with weight decay; proves GD converges to a neighbourhood of the global minimizers of the empirical loss, then partitions the input space into data-centered cells and decomposes population risk into three explicit terms: E_opt (optimization error on training data), E_PV (propagation-variation error: how predictions oscillate within each cell), and E_A (approximation error: how far ground truth is from the function class). The culprit behind memorization without generalization is E_PV: E_opt vanishes early while E_PV stays large. Under approximate Euler homogeneity, layerwise variation satisfies a differential inequality with explicit contraction rate lambda_l >= 1 - 2(1 + L) eta_2 lambda-bar, bounding E_PV by the initial hypothesis variation decaying exponentially in steps at a rate proportional to weight decay. Yields a closed-form grokking residency time (Steps ~ Omega(log(Initial_Var / budget) / (lambda-bar * eta_2 * sum L_l))): weight decay accelerates generalization, deeper networks contract faster, smoother initialization shortens the delay. Bounds are qualitative -- constants depend on suprema over parameter balls and loss smoothness, so they give rates and scaling laws, not exact step counts.
Keywords
generalization theory · grokking · weight decay · optimization · memorization
Topics
generalization theory, grokking, weight decay, optimization
Research notes
- Discovery: Grigory Sapunov (@che_shr_cat, verified) 2026-09-30 thread: https://x.com/che_shr_cat/status/2105179516738175168
- Thread takeaways: grokking is not magic and flat minima do not explain it -- weight decay is the active engine contracting the hypothesis between data points, driving memorization-to-generalization transition; delayed generalization is the signature of a function being smoothed by its own training dynamics.
- Sapunov's full technical breakdown: https://arxiviq.substack.com/p/a-theoretica...
- Connects to the collection's generalization-theory, grokking, and optimization entries.