← Back to explorer

How Model Growth, Recursion, and Boundary Operators Influence Scaling Exponents

Type
paper
Venue
arXiv:2609.19107
Year
1910
Source
arxiv
Access
public
Language
en
Added
2026-09-29
Verified
2026-09-29

Summary

Chen, Vegesna, Dahal & Wilson show that architecture — specifically model growth via looped/recurrent transformers — can modify pretraining scaling exponents, not just shift loss curves. A 7.4B growth architecture matches GPT-3 13B on CORE with roughly 20× less compute, with gains that increase with scale; even a simple boundary operator in a vanilla transformer helps.

Keywords

scaling-laws · looped-transformers · architecture · pretraining

Topics

scaling-laws, looped-transformers, architecture, pretraining

Research notes

  • Method: Prelude-core-coda transformer family with Loop-Grow (one shared core applied recurrently, loop count increased during training), an untied-growth control (distinct weights per added pass), and a boundary-operator variant; compares compute-optimal scaling ladders at matched FLOPs.
  • Key findings: Architectural interventions can change scaling exponents in pretraining — contrary to the conventional view that exponents are fixed; 7.4B model-growth architecture matches GPT-3 13B on CORE with ~20× less compute; efficiency gains grow with scale; Tied Loop-Grow: 1.36× compute multiplier over vanilla at 1e20 FLOPs; untied growth: 1.55× (independent briefing); A simple boundary operator (normalize + inject an earlier block) in a vanilla transformer also improves the exponent, to a lesser extent; In data-constrained multi-epoch settings, standard looping regularizes and it becomes compute-optimal to increase loop count with scale
  • Limitations: The ~20× GPT-3 comparison uses different data/eval pipelines — indicative, not a controlled gain. 44-page paper; independent replication pending.
  • 44 pages. v2 revised Sep 17, 2026. Framed through 'computational depth': for a fixed budget, maximize usable transformer depth.