← Back to explorer

Ceiling of a Task: When Can a Transformer Succeed Without Its Chain of Thought?

Type
paper
Venue
arXiv:2609.33134 (cs.AI), submitted 27 Sep 2026
Year
2026
Source
arxiv
Access
free
Language
English
Added
2026-10-01
Verified
2026-10-01

Summary

Asks whether chains of thought do real computational work or are decorative, by viewing a transformer as a shallow circuit: one forward pass has constant depth, so any constant number of passes is a shallow circuit, and the best accuracy such a circuit can reach on a task is the task's ceiling. Three results hold for every transformer on serial tasks (ceiling below one): Necessity — replacing the chain with content-independent filler or a restated question drops accuracy to the ceiling, and to chance on a maximally serial task; Depth — no shallow computation can write the chain of a model that exceeds the ceiling; Locality — the answer is one shallow pass from the finished chain, so all serial reasoning happens in the chain. Experiments confirm the predictions: on finite-group word problems, chain-trained small transformers solve every input length and fall to chance when the chain is erased, and on MATH-500/AIME erasing the chain costs open reasoning models 0.52 to 0.82 accuracy while a sentence shuffle is harmless and a token shuffle is as harmful as erasing.

Keywords

chain of thought · ceiling of a task · serial tasks · shallow circuits · transformer theory · filler tokens · MATH-500 · AIME · GRPO

Topics

chain of thought, reasoning models, circuit complexity, serial computation, transformer theory

Research notes

  • Discovery: @arXivBangers X post 2026-10-01 (https://x.com/arXivBangers/status/2105799962143666391)
  • Three theorems on serial tasks, holding for every transformer regardless of training: Necessity (content-independent chain replacements such as filler tokens or question restatement drive accuracy down to the ceiling, and to chance on a maximally serial task), Depth (no shallow computation can write, even approximately, the chain of a model whose accuracy exceeds the ceiling), Locality (the answer is one shallow pass away from the finished chain, so all serial reasoning happens in the chain)
  • Finite-group word problems (ceilings known): small transformers trained from scratch, with or without RL, attain the predicted numbers — chain-trained models solve every input length and fall to chance when the chain is erased; chainless models collapse to the ceiling as input length grows; open-weight reasoning models on the same problem in words return to baseline without their chain
  • MATH-500 and AIME: erasing the chain costs open reasoning models 0.52-0.82 accuracy; sentence shuffle harmless; token shuffle as harmful as erasing; same holds for GRPO checkpoints trained with a correct or a random reward
  • Affiliations (paper first page): Jiashu He and Alejandro Ribeiro, University of Pennsylvania; Jinxuan Fan and Radu Marculescu, University of Texas at Austin; Xiao Xiao, Yale University
  • No code release mentioned
  • License: CC BY 4.0