How Model Growth, Recursion, and Boundary Operators Influence Scaling Exponents
- Type
- paper
- Venue
- arXiv:2609.19107
- Year
- 1910
- Source
- arxiv
- Access
- public
- Language
- en
- Added
- 2026-09-29
- Verified
- 2026-09-29
Summary
Chen, Vegesna, Dahal & Wilson show that architecture — specifically model growth via looped/recurrent transformers — can modify pretraining scaling exponents, not just shift loss curves. A 7.4B growth architecture matches GPT-3 13B on CORE with roughly 20× less compute, with gains that increase with scale; even a simple boundary operator in a vanilla transformer helps.
Keywords
scaling-laws · looped-transformers · architecture · pretraining
Topics
scaling-laws, looped-transformers, architecture, pretraining
Research notes
- Method: Prelude-core-coda transformer family with Loop-Grow (one shared core applied recurrently, loop count increased during training), an untied-growth control (distinct weights per added pass), and a boundary-operator variant; compares compute-optimal scaling ladders at matched FLOPs.
- Key findings: Architectural interventions can change scaling exponents in pretraining — contrary to the conventional view that exponents are fixed; 7.4B model-growth architecture matches GPT-3 13B on CORE with ~20× less compute; efficiency gains grow with scale; Tied Loop-Grow: 1.36× compute multiplier over vanilla at 1e20 FLOPs; untied growth: 1.55× (independent briefing); A simple boundary operator (normalize + inject an earlier block) in a vanilla transformer also improves the exponent, to a lesser extent; In data-constrained multi-epoch settings, standard looping regularizes and it becomes compute-optimal to increase loop count with scale
- Limitations: The ~20× GPT-3 comparison uses different data/eval pipelines — indicative, not a controlled gain. 44-page paper; independent replication pending.
- 44 pages. v2 revised Sep 17, 2026. Framed through 'computational depth': for a fixed budget, maximize usable transformer depth.