Thought Branches: Interpreting LLM Reasoning Requires Resampling
- Type
- paper
- Venue
- arXiv
- Year
- 2026
- Source
- arxiv
- Access
- public
- Language
- en
- Added
- 2026-09-29
- Verified
- 2026-09-29
Summary
Argues that interpreting a single chain of thought is inadequate for understanding LLM reasoning; proposes resampling-based causal analysis of reasoning steps and a resilience metric for reasoning-step removal.
Keywords
interpretability · chain-of-thought · reasoning · mechanistic
Topics
interpretability, chain-of-thought, reasoning, mechanistic
Research notes
- Discovery: Shared in #random-papers as an arXiv link.
- Method: Downstream resampling to estimate the causal influence of individual reasoning steps; resilience metric measuring how reasoning degrades when steps are removed.
- Key findings: Self-preservation sentences had little causal influence in agentic-misalignment blackmail cases when tested by resampling, illustrating that a single CoT can mislead interpretability.
- Submitted 2025-10-31, revised 2026-04-13.