How Likely Do LLMs with CoT Mimic Human Reasoning?
- Type
- other
- Venue
- arXiv / Westlake University / Zhejiang University
Summary
Causal analysis of instruction Z, CoT X, and answer Y yields four SCM types (chain / common-cause / full-connection / isolation). GPT-3.5-Turbo is type I on GSM8K/FOLIO but type II on Addition (CoT explains rather than reasons). Across Llama2-7B/70B, GPT-3.5, and GPT-4, type III is most common (10/24 LLM-tasks); larger models do not reliably approach type I. ICL (2–8 shot) strengthens CoT→answer and weakens instruction→answer; SFT and DPO on Mistral-7B weaken the structure. Addition consistency errors 64.8% GPT-3.5 / 74.4% GPT-4 (incorrect CoT, correct answer). COLING 2025. Code https://github.com/StevenZHB/CoT_Causal_Analysis.
Keywords
cot · causal-analysis · scm · faithfulness · consistency · icl · sft · rlhf · coling · westlake
Topics
chain-of-thought, causal analysis, LLM reasoning
Research notes
- Primary: arxiv abs (cs.CL; also cs.AI, cs.LG). License CC BY 4.0 on HTML at check. Equal contrib Bao/Zhang; corresponding Yue Zhang. Westlake University / Zhejiang University; emails {baoguangsheng, zhanghongbo, wangcunxiang, yanglinyi, zhangyue}@westlake.edu.cn. Code on HTML footnote https://github.com/StevenZHB/CoT_Causal_Analysis (23 stars at check; no LICENSE file at check). HF paper page 0 upvotes; no linked models/datasets. Discord posted HTML v2. Eval uses public GSM8K/ProofWriter/FOLIO/LogiQA plus synthetic n-digit arithmetic, not a new hosted corpus, so no datasets_local row.