Understanding and Mitigating Premature Confidence for Better LLM Reasoning
- Type
- other
- Venue
- arXiv / Carnegie Mellon University
Summary
Defines premature confidence by probing truncated CoTs: the answer is often already fixed, so remaining tokens cannot causally shape it. On CSQA, prematurely confident traces have 2.8x more logical flaws than progressive ones (pattern holds on GPQA/LSAT/MuSR and on correct answers). Progressive confidence shaping subtracts η⟨c,w⟩ from GRPO advantages with a fixed decreasing vector w, no PRM labels. Hard Countdown Pass@1 19.1%→61.1% (+42.0pp) and issue rate 93.5%→45.5%; AIME Pass@64 +6.6pp; SciQA +2.9–5.8pp from 1.7B–8B. Also raises hint-acknowledgement on a safety benchmark. Premature confidence grows with model size and task difficulty (accessibility dominates utility). No official code on abs.
Keywords
premature-confidence · grpo · cot-faithfulness · process-reward · cmu · countdown · dapo · aime
Topics
LLM reasoning, chain-of-thought, RL, faithfulness
Research notes
- Primary: arxiv abs (cs.AI). CC BY 4.0 on abs HTML. CMU; correspondence via listed authors. Raghunathan NSF/UK AISI; Risteski NSF/Amazon/ONR/Google/OpenAI. No official code on abs. HF paper page 0 upvotes. Discord posted abs. Uses public CSQA/GPQA/LSAT/MuSR/DAPO/AIME plus synthetic Countdown; no new standalone public corpus, so no datasets_local row.