← Back to explorer

Telescopic Language Models

Type
paper
Venue
arXiv
Year
2026
Source
arxiv
Access
public
Language
en
Added
2026-09-29
Verified
2026-09-29

Summary

Trains a nested-capacity Transformer that is usable at many compute budgets: stochastic prefix supervision over layer prefixes plus a full-capacity anchor. A 200M proxy trained on 20B FineWeb-Edu tokens was usable at every one of 20 layer prefixes.

Keywords

efficient inference · early exit · nested models · pretraining

Topics

efficient inference, early exit, nested models, pretraining

Research notes

  • Discovery: Shared in #random-papers as an arXiv link.
  • Method: Nested-capacity Transformer; stochastic prefix supervision combined with a full-capacity anchor objective.
  • Key findings: Reported 43-44% better area under the quality-budget curve and ~12% lower GPU cost per run than fixed-exit model suites.
  • Full author list and exact paper title not fully captured; verify on arXiv before publishing.