Not All LLM Reasoning is Visible in the Chain-of-Thought
- Type
- paper
- Venue
- arXiv / NYU / UMD / TogetherAI
- Year
- 2026
- Source
- arxiv
- Access
- free
- Language
- en
- Added
- 2026-08-14T16:36:00Z
- Verified
- 2026-08-14T16:36:00Z
Summary
Defines invisible reasoning via filler-token diagnostics: accuracy rises with fixed, question-independent filler spans, depends on filler type, and the type ranking differs across models. Evaluates 13 frontier models on 4-digit multiplication, nested arithmetic, and variable-counting. Many models gain (up to +13 pp; Gemini 3 Flash +10.7 arithmetic; Opus 4.5 +11.2 arithmetic / +10.0 multiplication with counting fillers). Filler tokens let Opus 4.5 satisfy a hidden modular constraint (x mod 2 = 1: 33.5%→44.5% N/A) without hurting the primary task. RL on Qwen3-235B creates filler-type preferences but the test-time filler benefit does not persist; SFT also fails to transfer it.
Keywords
cot · filler-tokens · invisible-reasoning · ai-safety · qwen3 · claude · monitoring
Topics
AI safety, chain-of-thought, latent reasoning
Research notes
- Primary: arxiv abs (default nonexclusive-distrib 1.0, cs.CL/AI/LG). Affiliations NYU, University of Maryland, TogetherAI. No official code on abs. Discord posted abs link with a Substack UTM query.