Beyond Binary Rewards: Training LMs to Reason About Their Uncertainty
- Type
- paper
- Venue
- arXiv
- Year
- 2026
- Source
- arxiv
- Access
- public
- Language
- en
- Added
- 2026-09-30
- Verified
- 2026-09-30
Summary
Almost all successful RL for reasoning uses binary rewards that evaluate output correctness; because they do not penalize guessing or low-confidence outputs, they degrade calibration and increase hallucination in other domains. RLCR (Reinforcement Learning with Calibration Rewards) trains reasoning models to generate predictions and numerical confidence estimates after reasoning, optimizing a reward that augments a binary correctness score with a Brier score -- a scoring rule for confidence estimates that incentivizes calibrated prediction. The paper proves that this reward (or any bounded proper scoring rule) yields models that are both accurate and well-calibrated, shows across diverse datasets that RLCR substantially improves calibration with no loss in accuracy in-domain and out-of-domain (outperforming ordinary RL and post-hoc confidence classifiers), and demonstrates that verbalized confidence can be leveraged at test time via confidence-weighted scaling. Code, models, and info at rl-calibration.github.io.
Keywords
reinforcement learning · calibration · reasoning · uncertainty · RLCR
Topics
calibration, reinforcement learning, reasoning, uncertainty
Research notes
- Discovery: linked from Sebastian Raschka's 2026-09-29 essay "Language Models for Text Classification: From Bag-of-Words to Jev" (https://magazine.sebastianraschka.com/p/classifier-history-and-jev) as the public related method (RLCR) to TypeSafe AI's proprietary calibration training for Jev; HotpotQA ECE 0.37 to 0.03 vs RLVR.
- Submitted 2025-07-22, revised 2026-05-15, cs.LG.
- Connects to the collection's RL, calibration, and LLM-judge entries.