Telescopic Language Models
- Type
- paper
- Venue
- arXiv
- Year
- 2026
- Source
- arxiv
- Access
- public
- Language
- en
- Added
- 2026-09-29
- Verified
- 2026-09-29
Summary
Trains a nested-capacity Transformer that is usable at many compute budgets: stochastic prefix supervision over layer prefixes plus a full-capacity anchor. A 200M proxy trained on 20B FineWeb-Edu tokens was usable at every one of 20 layer prefixes.
Keywords
efficient inference · early exit · nested models · pretraining
Topics
efficient inference, early exit, nested models, pretraining
Research notes
- Discovery: Shared in #random-papers as an arXiv link.
- Method: Nested-capacity Transformer; stochastic prefix supervision combined with a full-capacity anchor objective.
- Key findings: Reported 43-44% better area under the quality-budget curve and ~12% lower GPU cost per run than fixed-exit model suites.
- Full author list and exact paper title not fully captured; verify on arXiv before publishing.