DepthBench: Measuring How Residual Connections Enable More Computational Depth
- Type
- paper
- Venue
- arXiv
- Year
- 2026
- Source
- arxiv
- Access
- public
- Language
- en
- Added
- 2026-09-29
- Verified
- 2026-09-29
Summary
Introduces DepthBench, a controlled benchmark for studying computational depth across architectures, proposing effective computational depth as a new scaling axis. Systematically varies the width-depth aspect ratio (d_model/n_layer) from shallow-wide to deep-narrow shapes while keeping model size and the pretraining recipe fixed, across 10 representative architectures. Finds the benefit of allocating capacity to depth is strongly architecture-dependent: standard Pre-LN and most norm- and scaling-based variants (e.g. LayerNorm Scaling) give little benefit and can even degrade as models get deeper and narrower, whereas HC and Full AttnRes improve consistently even at extreme deep shapes. Gains extend beyond pretraining loss into improved domain-specific performance and effective computation; layer-level analyses link HC/Full AttnRes gains to more effective utilization of additional layers via distinct mechanisms. Concludes residual-connection design determines whether depth can serve as a meaningful scaling axis.
Keywords
architecture · scaling laws · depth · residual connections · mHC · AttnRes · benchmark · transformers
Topics
architecture, scaling, residual connections, depth
Research notes
- Discovery: commented on by Teortaxes (@teortaxesTex) on 2026-09-29 in a quote-tweet of the paper announcement by Shiwei Liu (@Shiwei_Liu66): https://x.com/teortaxesTex/status/2105090036530086354
- Teortaxes notes: in an experimental 400M model, mHC saturates at 24 layers, but is likely at least no worse than AttnRes in the wide-and-shallow regime targeted by DeepSeek V4.1-Flash; would like to see internal lab ablations.
- Submitted to arXiv 2026-09-26, cs.CV/cs.AI/cs.LG.
- Connects to the collection's residual-connection entries (mHC, AttnRes, KEEL, LNS, MoDA) and to width-depth aspect-ratio / depth-scaling research. author thread link https://x.com/Keyuciallo/status/2105057936166801585 (Keyu Wang, 2026-09-30), code https://github.com/keyu-wang-2002/DepthBench, models https://huggingface.co/aspect-ratio-scaling.