← Back to explorer

Thought Branches: Interpreting LLM Reasoning Requires Resampling

Type
paper
Venue
arXiv
Year
2026
Source
arxiv
Access
public
Language
en
Added
2026-09-29
Verified
2026-09-29

Summary

Argues that interpreting a single chain of thought is inadequate for understanding LLM reasoning; proposes resampling-based causal analysis of reasoning steps and a resilience metric for reasoning-step removal.

Keywords

interpretability · chain-of-thought · reasoning · mechanistic

Topics

interpretability, chain-of-thought, reasoning, mechanistic

Research notes

  • Discovery: Shared in #random-papers as an arXiv link.
  • Method: Downstream resampling to estimate the causal influence of individual reasoning steps; resilience metric measuring how reasoning degrades when steps are removed.
  • Key findings: Self-preservation sentences had little causal influence in agentic-misalignment blackmail cases when tested by resampling, illustrating that a single CoT can mislead interpretability.
  • Submitted 2025-10-31, revised 2026-04-13.