← Back to explorer

Not All LLM Reasoning is Visible in the Chain-of-Thought

Type
paper
Venue
arXiv / NYU / UMD / TogetherAI
Year
2026
Source
arxiv
Access
free
Language
en
Added
2026-08-14T16:36:00Z
Verified
2026-08-14T16:36:00Z

Summary

Defines invisible reasoning via filler-token diagnostics: accuracy rises with fixed, question-independent filler spans, depends on filler type, and the type ranking differs across models. Evaluates 13 frontier models on 4-digit multiplication, nested arithmetic, and variable-counting. Many models gain (up to +13 pp; Gemini 3 Flash +10.7 arithmetic; Opus 4.5 +11.2 arithmetic / +10.0 multiplication with counting fillers). Filler tokens let Opus 4.5 satisfy a hidden modular constraint (x mod 2 = 1: 33.5%→44.5% N/A) without hurting the primary task. RL on Qwen3-235B creates filler-type preferences but the test-time filler benefit does not persist; SFT also fails to transfer it.

Keywords

cot · filler-tokens · invisible-reasoning · ai-safety · qwen3 · claude · monitoring

Topics

AI safety, chain-of-thought, latent reasoning

Research notes

  • Primary: arxiv abs (default nonexclusive-distrib 1.0, cs.CL/AI/LG). Affiliations NYU, University of Maryland, TogetherAI. No official code on abs. Discord posted abs link with a Substack UTM query.