cs.CLSep 27, 2026

Calibration, Not Answer Selection: Distilling Internal Confidence in Reasoning Models

Authors: Yadong Xi, Rongsheng Zhang, Tangjie Lv, Ziyang Luo, Ruochen Zhao

Organizations: NetEase · Amazon · Singapore University of Technology and Design

Abstract

Reinforcement learning with binary correctness rewards trains correctness, not calibrated confidence. The confidence that reasoning models verbalize is systematically overconfident, and the problem is not merely one of scale: verbalized confidence tracks how willing a model is to commit to an answer, not how likely the answer is to be right. Post-hoc rescaling therefore fits one distribution but rarely transfers. We look inside the model instead. On factual question answering, a linear probe on the hidden state between the chain of thought and the answer is substantially better calibrated: its expected calibration error is 5 to 38 times lower than that of the verbalized score across four benchmarks and two model families. However, when used to pick among N sampled answers, that same probe nearly ties majority voting yet falls far short of the oracle. Internal states answer "how certain am I" well and "which answer is right" poorly, so the signal should be reported as a confidence rather than used to select answers. As a result, we introduce probe-guided self-distillation (Probe-SD): score a model's own sampled traces with the probe, overwrite the confidence each trace states, and finetune the base checkpoint of the same family, so nothing but the model itself remains at test time. On Qwen3-14B, Probe-SD cuts ECE from 0.178 to 0.024 in-domain and from 0.542 to 0.113 out-of-domain, where it also beats post-hoc recalibration and self-consistency distillation. The resulting confidence is well-calibrated and useful for weighted voting, behaviors previously attributed to online RL, here obtained with supervised finetuning alone.

Figures & tables

Appendix figures & tables27 assets

Supplementary material from the paper’s appendix.

Appendix

Explore similar work

CardsList
  1. Beyond Verbalized Confidence: Calibrating Reasoners with Differentiable Readouts

    Sep 28, 2026Chenxiao Fan, Chongming Gao, Gangyi Zhang +8Reinforcement Learning With Verifiable RewardVerifiable Rewards

  2. Post-Training Shifts Confidence: A Three-Stage Analysis of How SFT, RL, and OPD Shape Pre-, Intra-, and Post-CoT Calibration

    Jul 15, 2026Shuhao Li, Guodong Du, Anhao Zhao +3Post-TrainingSupervised Finetuning