Calibration, Not Answer Selection: Distilling Internal Confidence in Reasoning Models
Organizations: NetEase · Amazon · Singapore University of Technology and Design
Abstract
Reinforcement learning with binary correctness rewards trains correctness, not calibrated confidence. The confidence that reasoning models verbalize is systematically overconfident, and the problem is not merely one of scale: verbalized confidence tracks how willing a model is to commit to an answer, not how likely the answer is to be right. Post-hoc rescaling therefore fits one distribution but rarely transfers. We look inside the model instead. On factual question answering, a linear probe on the hidden state between the chain of thought and the answer is substantially better calibrated: its expected calibration error is 5 to 38 times lower than that of the verbalized score across four benchmarks and two model families. However, when used to pick among N sampled answers, that same probe nearly ties majority voting yet falls far short of the oracle. Internal states answer "how certain am I" well and "which answer is right" poorly, so the signal should be reported as a confidence rather than used to select answers. As a result, we introduce probe-guided self-distillation (Probe-SD): score a model's own sampled traces with the probe, overwrite the confidence each trace states, and finetune the base checkpoint of the same family, so nothing but the model itself remains at test time. On Qwen3-14B, Probe-SD cuts ECE from 0.178 to 0.024 in-domain and from 0.542 to 0.113 out-of-domain, where it also beats post-hoc recalibration and self-consistency distillation. The resulting confidence is well-calibrated and useful for weighted voting, behaviors previously attributed to online RL, here obtained with supervised finetuning alone.
Figures & tables
| Qwen3-14B | Ministral-3-8B | ||||||||||
| Dataset | Acc | ECE | AUROC | Brier | Cov. | Acc | ECE | AUROC | Brier | Cov. | |
| TriviaQA (ID) | V | .729 | .178 .020 | .825 .017 | .184 .017 | .980 | .664 | .247 .024 | .669 .015 | .263 .021 | .957 |
| P | .015 .009 | .948 .009 | .081 .008 | 1.000 | .014 .008 | .917 .012 | .105 .009 | .977 | |||
| EntityQ | V | .267 | .529 .022 | .812 .017 | .455 .018 | .903 | .207 | .627 .022 | .734 .016 | .564 .019 | .950 |
| P | .014 .009 | .913 .014 | .097 .010 | 1.000 | .024 .012 | .909 .015 | .091 .009 | .999 | |||
| NQ-open | V | .440 | .436 .029 | .705 .021 | .416 .025 | .723 | .370 | .513 .026 | .643 .018 | .490 .022 | .798 |
| In-domain (TriviaQA) | Out-of-domain (macro of 3) | ||||||||
| Model | Method | Acc | ECE | AUROC | Brier | Acc | ECE | AUROC | Brier |
| Qwen3-14B | Verbal | .729 .024 | .178 .020 | .825 .017 | .184 .017 | .258 .013 | .542 .013 | .715 .017 | .468 .011 |
| Verbal-Iso | — | .013 .009 | .825 .017 | .128 .010 | — | .274 .012 | .714 .017 | .236 .007 | |
| Verbal-TS+bias | — | .047 .011 | .825 .017 | .131 .010 | — | .294 .011 | .715 .017 | .245 .009 | |
| Answer Probability | — | .251 .024 | .636 .029 | .253 .023 | — | .695 .013 | .586 .020 | .686 .013 | |
| Surface-Feat | — | .042 .010 | .858 .016 | .128 .010 | — | .332 .012 | .697 .017 | .285 .007 | |
| Components | In-domain | Out-of-domain | ||||||||
| Variant | relabel | CW | Acc | ECE | AUROC | Brier | Acc | ECE | AUROC | Brier |
| Probe-SD (full) | yes | yes | .749 | .024 | .928 | .092 | .279 | .113 | .783 | .147 |
| w/o CW | yes | no | .748 | .033 | .919 | .097 | .280 | .129 | .768 | .157 |
| w/o relabel | no | no | .747 | .159 | .814 | .173 | .281 | .522 | .723 | .446 |
Appendix figures & tables27 assets
Supplementary material from the paper’s appendix.
Appendix
| qid | idx | Gold | Model answer | Graded | Verbal new | |
| qb_8463 | 10 | Times Square | Times Square | correct | 95 73 | .727 |
| qw_10780 | 17 | USA | United States of America | correct | 95 74 | .742 |
| jp_2478 | 17 | Elizabeth | Elizabeth | correct | 99 81 | .806 |
| sfq_19580 | 11 | Chris Huhne | Chris Huhne | correct | 95 75 | .749 |
| sfq_4960 | 0 | Glasgow | York | incorrect | 85 66 | .665 |
| Dimension | Qwen3-14B | Ministral-3-8B |
| Layers | 40 | 34 |
| Regularization | ||
| Readout | ||
| Total configurations | 1,760 | 1,496 |
| Dataset | Readout | sel. Acc | ECE | AUROC | Brier |
| Qwen3-14B | |||||
| TriviaQA | (i) residual | .772 | .015 | .948 | .081 |
| (ii) MLP | .779 | .019 | .948 | .081 | |
| EntityQ | (i) residual | .310 | .014 | .913 | .097 |
| (ii) MLP | .302 | .026 | .909 | .100 | |
| NQ-open | (i) residual | .463 | .078 | .826 | .177 |
| Model | Dataset | Oracle (pass@ ) | Max-Probe | Majority vote |
| Qwen3-14B | TriviaQA | .879 | .778 | .771 |
| EntityQ | .447 | .315 | .306 | |
| NQ-open | .695 | .458 | .492 | |
| SimpleQA | .246 | .074 | .077 | |
| Ministral-3-8B | TriviaQA | .844 | .717 | .728 |
| EntityQ | .363 | .231 | .237 |
| Metric | String pre-screen | After human review |
| Extraction accuracy | 92.0% (184/200) | 99.5% (199/200) |
| Grading accuracy (answerable) | 93.8% (181/193) | 99.5% (198/199) |
| Grading precision | 83.1% | 98.7% (77/78) |
| Grading recall | 100% | 100% (77/77) |
| True grading errors | — | 1/200 (0.5%) |
| Extraction outcome | Verdict | Flagged grading case | Verdict | ||
| Exact or prefix match | 175 | faithful | More specific answer contains gold | 5 | correct |
| Substring match | 9 | correct distillation | Alias not in gold list | 3 | correct |
| Surface mismatch, valid | 5 | valid synthesis | Functional description entity | 2 | correct |
| Empty: explicit abstention | 10 | correct empty | Output contains gold substring | 3 | correct |
| Empty: answer was given | 1 | true error | Subset or containment relation | 4 | correct |
| Diacritic difference | 1 | correct |
| Source | ECE | AUROC | Brier | |||
| Full | Subset | Full | Subset | Full | Subset | |
| Probe | .0147 | .0363 | .9477 | .9621 | .0809 | .0718 |
| Verbal | .1782 | .1905 | .8243 | .8511 | .1839 | .1877 |
| Token | .2514 | .2602 | .6360 | .6181 | .2530 | .2586 |
| Dataset (Acc) | Source | Brier | Rel | Res | Unc |
| Qwen3-14B | |||||
| TriviaQA (.729) | Verbal | .184 | .052 | .059 | .194 |
| Probe | .081 | .001 | .116 | .197 | |
| EntityQ (.267) | Verbal | .455 | .297 | .040 | .207 |
| Probe | .097 | .000 | .098 | .196 | |
| NQ-open (.440) | Verbal | .416 | .207 | .023 | .249 |
| Model | Method | TriviaQA | EntityQ | NQ-open | SimpleQA | Macro |
| ECE | ||||||
| Qwen | Probe-SD | .023 .006 | .090 .008 | .124 .012 | .114 .007 | .088 .004 |
| Qwen | Verbal ∗ | .177 .021 | .531 .022 | .423 .032 | .665 .015 | .449 .012 |
| Qwen | Probe-SD-LRM | .030 .007 | .102 .008 | .163 .013 | .136 .007 | .108 .004 |
| Qwen | Probe-SD w/o CW | .032 .006 | .105 .008 | .146 .012 | .126 .007 | .102 .004 |
| Qwen | SC-SD | .063 .008 | .224 .009 | .241 .013 | .237 .009 | .191 .005 |
| Model | Method | TriviaQA | EntityQ | NQ-open | SimpleQA | Macro |
| ECE | ||||||
| Qwen | Probe-SD | .024 .009 | .089 .013 | .162 .018 | .104 .010 | .095 .006 |
| Qwen | Verbal ∗ | .182 .020 | .493 .020 | .410 .022 | .545 .018 | .408 .010 |
| Qwen | Probe-SD-LRM | .031 .010 | .101 .014 | .205 .018 | .122 .010 | .115 .007 |
| Qwen | Probe-SD w/o CW | .033 .010 | .105 .013 | .182 .017 | .114 .010 | .109 .006 |
| Qwen | SC-SD | .064 .013 | .222 .016 | .266 .019 | .208 .012 | .190 .007 |
| Dataset | Convention | ECE | AUROC | Brier |
| TriviaQA | scorable subset | .014 [.010,.026] | .948 [.938,.957] | .080 [.072,.089] |
| common subset ( ) | .013 [.010,.025] | .948 [.938,.957] | .080 [.071,.089] | |
| full-sample attribution | .015 [.010,.027] | — | .080 [.072,.088] | |
| EntityQ | scorable subset | .022 [.015,.038] | .898 [.882,.914] | .109 [.099,.120] |
| common subset ( ) | .023 [.015,.039] | .899 [.883,.915] | .109 [.099,.120] | |
| full-sample attribution | .020 [.014,.035] | — | .099 [.090,.109] |
| Hyperparameter | Qwen3-14B | Ministral-3-8B | Note |
| Training set size | 10,000 | 10,000 | identical across methods |
| Learning rate | |||
| Epochs | 2 | 2 | steps |
| LR schedule | cosine to 0 | cosine to 0 | 50 warmup steps |
| Weight decay | 0.1 | 0.1 | |
| Adam | 0.9 / 0.95 | 0.9 / 0.95 |
| Purpose | Qwen3-14B | Ministral-3-8B | |
| Evaluation probe scoring | 0.6 | L24 residual, | L21 MLP, |
| Probe-SD training labels | 1.5 | L24 residual, | L22 residual, |
| Training pool | TriviaQA | NQ-open | EntityQ | SimpleQA |
| .739 .024 | .470 .027 | .289 .025 | .077 .013 | |
| .738 .024 | .467 .026 | .288 .025 | .078 .013 | |
| .744 .024 | .470 .026 | .291 .025 | .076 .012 | |
| , no top- /top- | .745 .024 | .471 .026 | .286 .024 | .078 .012 |
| Dataset | Method | Acc | ECE | AUROC | Brier |
| TriviaQA | Verbal | .729 .024 ∗ | .178 .020 | .825 .017 | .184 .017 |
| Frozen Probe ∗ | — | .014 .008 | .948 .010 | .080 .009 | |
| Verbal-Iso | — | .013 .009 | .825 .017 | .128 .010 | |
| Verbal-TS+bias | — | .047 .011 | .825 .017 | .131 .010 | |
| Answer Prob. | — | .251 .024 | .636 .029 | .253 .023 | |
| Surface-Feat | — | .042 .010 | .858 .016 | .128 .010 |
| Dataset | Method | Acc | ECE | AUROC | Brier |
| TriviaQA | Verbal | .664 .025 ∗ | .247 .024 | .669 .015 | .263 .021 |
| Frozen Probe ∗ | — | .019 .008 | .915 .013 | .106 .009 | |
| Verbal-Iso | — | .020 .014 | .669 .015 | .188 .009 | |
| Verbal-TS+bias | — | .067 .016 | .669 .015 | .193 .008 | |
| Answer Prob. | — | .302 .025 | .569 .027 | .306 .025 | |
| Surface-Feat | — | .042 .010 | .770 .016 | .174 .009 |
| Model | Set | mean | p25 | p50 | p75 | |
| Ministral-3-8B | Pool (random 15k) | 15,000 | 0.636 | 0.317 | 0.777 | 0.943 |
| Retained | 13,840 | 0.672 | 0.495 | 0.782 | 0.930 | |
| Qwen3-14B | Pool (random 15k) | 15,000 | 0.722 | 0.460 | 0.936 | 0.992 |
| Retained | 14,951 | 0.723 | 0.485 | 0.915 | 0.989 |
| Model | Downsampling | ECE | AUROC | Brier | Sample Acc | pp | Train Acc |
| Qwen3-14B | calibrated (Probe-SD) | 0.091 | 0.819 | 0.134 | 0.397 | — | — |
| random ( rand_sub ) | 0.094 | 0.820 | 0.134 | 0.394 | — | — | |
| Ministral-3-8B | random | 0.123 | 0.808 | 0.150 | 0.346 | 4.26 | 70.0% |
| diffdist_rp | 0.112 | 0.813 | 0.144 | 0.347 | 0.01 | 65.5% | |
| calibrated | 0.105 | 0.815 | 0.142 | 0.348 | 1.17 | 61.8% | |
| probe_align | 0.118 | 0.807 | 0.149 | 0.348 | 0.01 | 66.2% |
| Decile | Probe | Observed acc | Probe-SD | SC-SD | Verbal-SD | |
| 1 | .675 | .191 | .141 | .295 | .374 | .756 |
| 2 | .803 | .337 | .331 | .424 | .522 | .817 |
| 3 | .866 | .461 | .423 | .537 | .637 | .867 |
| 4 | .907 | .601 | .546 | .645 | .752 | .906 |
| 5 | .938 | .836 | .891 | .865 | .878 | .936 |
| 6 | .952 | .900 | .916 | .904 | .926 | .945 |
| Quartile | Gap range | Probe-SD | SC-SD | Verbal-SD | itself |
| Q0 | .018 | .034 | .037 | .030 | |
| Q1 | .025 | .030 | .019 | .013 | |
| Q2 | .120 | .134 | .202 | .188 | |
| Q3 | .164 | .263 | .588 | .575 |
| Probe-SD | SC-SD | Verbal-SD | ||||
| Dataset | Probe | Probe | Probe | |||
| TriviaQA | .913 | .771 | .872 | .791 | .767 | .822 |
| EntityQ | .895 | .676 | .810 | .738 | .658 | .803 |
| NQ-open | .823 | .550 | .700 | .590 | .526 | .725 |
| SimpleQA | .762 | .494 | .642 | .641 | .468 | .757 |
| Bin | Probe | SC | Probe-SD | SC-SD | ||
| Probe SC | 23 | .635 | [ ] | [ ] | [ ] | [ ] |
| Probe SC | 86 | .700 | [ ] | [ ] | [ ] | [ ] |
| 579 | .739 | [ ] | [ ] | [ ] | [ ] | |
| Probe SC | 178 | .610 | [ ] | [ ] | [ ] | [ ] |
| Probe SC | 134 | .766 | [ ] | [ ] | [ ] | [ ] |
| Quintile | mean | Probe-SD | SC-SD | Verbal-SD | |
| Q0 | .030 | .055 | .115 | ||
| Q1 | .053 | .079 | .173 | ||
| Q2 | .123 | .171 | .327 | ||
| Q3 | .103 | .146 | .263 | ||
| Q4 | .100 | .126 | .179 |
| Qwen3-14B: SFT vs Instruct | Ministral-3-8B: SFT vs Reasoning | |
| Sample accuracy | .7508 vs .7272 ( ) | .6995 vs .6639 ( ) |
| Retrieval term | (93.7%) | (73.0%) |
| Usage term | (0.8%) | (19.9%) |
| Miss-branch term | (5.5%) | (7.0%) |
| Qwen3-14B | Ministral-3-8B | |||
| Metric | SFT-Base | Instruct | SFT-Base | Reasoning |
| Facts per trace | 9.24 | 9.02 | 9.33 | 9.29 |
| truth: true | 80.1% | 78.0% | 69.7% | 67.8% |
| truth: false (hallucination) | 13.5% | 15.2% | 26.7% | 27.6% |
| relevance: bridge | 5.83 | 5.97 | 6.13 | 5.57 |
| true bridge | 4.95 | 4.98 | 4.81 | 4.32 |
| Qwen3-14B | Ministral-3-8B | |||||
| Metric | SFT-Base | Instruct | SFT-Base | Reasoning | ||
| Gold recall hit rate | .7888 | .7640 | pp | .7294 | .6998 | pp |
| Accuracy gold recalled | .8915 | .8912 | pp | .8889 | .8791 | pp |
| Accuracy gold not recalled | .2254 | .1962 | pp | .1888 | .1622 | pp |
| Sample accuracy | .7508 | .7272 | pp | .6995 | .6639 | pp |
| Qwen3-14B | Ministral-3-8B | SFT-Base | ||||||
| Difficulty | SFT | Instr. | SFT | Reas. | Source side | easy | med. | hard |
| easy | 687 | 663 | 582 | 563 | Qwen Instruct easy (663) | 635 | 27 | 1 |
| medium | 179 | 190 | 255 | 245 | Qwen medium (190) | 51 | 130 | 9 |
| hard | 134 | 147 | 163 | 192 | Qwen hard (147) | 1 | 22 | 124 |
| Min. Reasoning easy (563) | 506 | 56 | 1 | |||||
| Min. medium (245) | 73 | 151 | 21 | |||||