Ceiling of a Task: When Can a Transformer Succeed Without Its Chain of Thought?
- Type
- paper
- Venue
- arXiv:2609.33134 (cs.AI), submitted 27 Sep 2026
- Year
- 2026
- Source
- arxiv
- Access
- free
- Language
- English
- Added
- 2026-10-01
- Verified
- 2026-10-01
Summary
Asks whether chains of thought do real computational work or are decorative, by viewing a transformer as a shallow circuit: one forward pass has constant depth, so any constant number of passes is a shallow circuit, and the best accuracy such a circuit can reach on a task is the task's ceiling. Three results hold for every transformer on serial tasks (ceiling below one): Necessity — replacing the chain with content-independent filler or a restated question drops accuracy to the ceiling, and to chance on a maximally serial task; Depth — no shallow computation can write the chain of a model that exceeds the ceiling; Locality — the answer is one shallow pass from the finished chain, so all serial reasoning happens in the chain. Experiments confirm the predictions: on finite-group word problems, chain-trained small transformers solve every input length and fall to chance when the chain is erased, and on MATH-500/AIME erasing the chain costs open reasoning models 0.52 to 0.82 accuracy while a sentence shuffle is harmless and a token shuffle is as harmful as erasing.
Keywords
chain of thought · ceiling of a task · serial tasks · shallow circuits · transformer theory · filler tokens · MATH-500 · AIME · GRPO
Topics
chain of thought, reasoning models, circuit complexity, serial computation, transformer theory
Research notes
- Discovery: @arXivBangers X post 2026-10-01 (https://x.com/arXivBangers/status/2105799962143666391)
- Three theorems on serial tasks, holding for every transformer regardless of training: Necessity (content-independent chain replacements such as filler tokens or question restatement drive accuracy down to the ceiling, and to chance on a maximally serial task), Depth (no shallow computation can write, even approximately, the chain of a model whose accuracy exceeds the ceiling), Locality (the answer is one shallow pass away from the finished chain, so all serial reasoning happens in the chain)
- Finite-group word problems (ceilings known): small transformers trained from scratch, with or without RL, attain the predicted numbers — chain-trained models solve every input length and fall to chance when the chain is erased; chainless models collapse to the ceiling as input length grows; open-weight reasoning models on the same problem in words return to baseline without their chain
- MATH-500 and AIME: erasing the chain costs open reasoning models 0.52-0.82 accuracy; sentence shuffle harmless; token shuffle as harmful as erasing; same holds for GRPO checkpoints trained with a correct or a random reward
- Affiliations (paper first page): Jiashu He and Alejandro Ribeiro, University of Pennsylvania; Jinxuan Fan and Radu Marculescu, University of Texas at Austin; Xiao Xiao, Yale University
- No code release mentioned
- License: CC BY 4.0