Reinforcement learning with verifiable rewards (RLVR) trains reasoning models to produce correct answers, but does not ensure that their stated confidence is calibrated. The resulting models are systematically overconfident. Recent methods train calibration inside the RLVR loop by having the model state a numerical confidence alongside its answer, but they all obtain the confidence by sampling it as text. This choice imposes two costs: a sampled confidence introduces variance and in practice collapses to a handful of distinct values, and sampling makes the confidence non-differentiable, forcing the calibration loss through a scalar reward. We propose CREDO (Confidence REaDOut) to replace sampling with a deterministic readout. While RLVR optimizes correctness, CREDO reads the confidence from a dedicated token pair in the model's output distribution and trains it by differentiable regression. CREDO further turns the trained confidence into a signal for accuracy, weighting rollouts by how far confidence and outcome disagree, so that accuracy and calibration improve together. Across mathematical and code reasoning, CREDO attains the best accuracy and calibration, and the gains extend to abstention and selective prediction.
Figures & tables
Figure 1: (a) The calibration task: a model answers and reports a confidence; calibration means the reported probability matches observed correctness. (b) Verbalized methods sample the confidence as text and train it through a scalar reward. CREDO reads the relative probability of a pair of reserved tokens and trains it by direct regression from the calibration loss. (c) Confidence distributions on code. Verbalized reports (GRPO, RLCR, DCPO) concentrate on a few values shared by correct and incorrect answers, even when calibration is trained explicitly (RLCR, DCPO). The readout (CREDO) spreads across the full range.
Figure 2: Collapse of the verbalized channel, with CREDO for contrast. Effective levels =1/∑vpv2 (equiprobable-level equivalent); top-value share = pool fraction of the most frequent value; tie tax = AUROC lost to exactly equal reports.
Figure 3: Overview of CREDO. Each rollout contains an answer, a confidence analysis, and a confidence slot. The readout c is the relative probability of two reserved tokens at the slot. The calibration loss trains c by differentiable regression toward a hybrid target t . The policy gradient scores the answer and analysis segments through their text. Discrepancy weighting gives more weight to rollouts where the outcome and the readout disagree.
Mathematics
Code
Method
Acc ↑
ECE ↓
AUROC ↑
Brier ↓
Acc ↑
ECE ↓
AUROC ↑
Brier ↓
Base
.439 ± .001
.472 ± .002
.644 ± .004
.453 ± .002
.455 ± .001
.418 ± .001
.690 ± .004
.395 ± .002
GRPO
.706 ± .009
.199 ± .027
.775 ± .062
.204 ± .029
.581 ± .007
.182 ± .070
.753 ± .061
.205 ± .066
RLCR
.682 ± .010
.159 ± .012
.831 ± .009
.171 ± .009
.553 ± .013
.075 ± .005
.815 ± .014
.146 ± .006
DCPO
.722 ± .008
.217 ± .008
.848 ± .012
.198 ± .003
.576 ± .003
.153 ± .024
.776 ± .025
.153 ± .024
Base-d
.439 ± .001
.442 ± .001
.765 ± .010
.424 ± .001
.455 ± .001
.389 ± .002
.782 ± .003
.373 ± .002
Table 1: Accuracy and calibration in and out of domain, each entry macro-averaged over evaluation sets and training seeds. Column groups are the evaluation domain. Baselines are read through their own verbalized numeral and again through a digit expectation (Appendix B.3 ), in the rows named -d ; CREDO is read through its readout. Bold marks the best mean in each column and any within one joint standard error of it.
Mathematics
Code
Acc ↑
ECE ↓
AUROC ↑
Brier ↓
Acc ↑
ECE ↓
AUROC ↑
Brier ↓
CREDO
.753
.094
.933
.088
.614
.052
.869
.117
w/o differentiable regression ( α=0 )
.696
.241
.880
.232
.611
.172
.857
.172
w/o uncertainty-analysis segment
.709
.096
.902
.107
.587
.048
.868
.117
w/o confidence-segment advantage ( A^conf≡0 )
.696
.099
.909
.102
.576
.047
.867
.126
w/o discrepancy weighting ( κ=0 )
.721
.122
.904
.102
.572
.051
.843
.132
Table 2: Leave-one-out ablation of CREDO (seed 43).
Figure 4: Accuracy and Brier for the two evaluation domains as each coefficient ( κ , α , γ ) of § 3 is varied on seed 43, the others held fixed. Exact values are in Table D.2 .
Figure 7
Appendix figures & tables13 assets
Supplementary material from the paper’s appendix.
Appendix
Shared
Base model
Qwen3-8B
Precision
bf16
Seeds
43 , 44 , 45
Rollouts per question G
8
Learning rate
2×10−6
Per-device batch size
2
Schedule
constant, no warmup
Gradient accumulation
64
Optimizer
AdamW ( Loshchilov and Hutter, 2019 )
Epochs
2
Rollout temperature
1.0
Rollout top- k
50
Appendix
Table 4: Training configuration. Values that differ by domain are given as mathematics / code. Rollouts are generated with vLLM ( Kwon et al., 2023 ) .
Mathematics
Code
Set
Problems
n
Set
Problems
n
DeepScaleR (train)
9,500
—
DeepCoder (train)
9,500
—
DeepScaleR test
500
4
DeepCoder test
500
4
MATH-500
500
4
HumanEval+
163
4
AIME 2024
30
8
LiveCodeBench v5
279
4
AIME 2025
30
8
LiveCodeBench v6
131
4
Appendix
Table 5: Datasets. n is the number of samples drawn per problem at evaluation. AIME and AMC are the corresponding years’ competition problems; HumanEval+ strengthens the tests of HumanEval ( Chen et al., 2021 ) .
Qwen3-1.7B
Qwen3-4B
Qwen3-8B
Acc ↑
ECE ↓
AUROC ↑
Brier ↓
Acc ↑
ECE ↓
AUROC ↑
Brier ↓
Acc ↑
ECE ↓
AUROC ↑
Brier ↓
Base
.268
.531
.691
.477
.425
.468
.641
.447
.440
.474
.640
.455
GRPO
.460
.391
.782
.357
.660
.221
.813
.220
.717
.174
.820
.179
RLCR
.494
.145
.853
.155
.663
.186
.836
.185
.682
.167
.837
.181
DCPO
.501
.203
.823
.189
.651
.220
.797
.196
.715
.222
.846
.195
Base-d
.268
.557
.717
.495
.425
.446
.765
.428
.440
.443
.772
.425
Appendix
Table 6: Mathematics results at three model scales (seed 43). All methods follow the same protocol as Table 1 . -d rows evaluate through the digit expectation; CREDO through its logit readout. Bold marks the best in each column.
Score-function reward
Differentiable regression
Verbalized numeral
RLCR / DCPO (Table 1 )
Digit-exp. emission
{H,L} readout
Readout + reward only
CREDO (Table 1 )
Appendix
Table 7: Design space for the confidence channel.
Mathematics
Code
Acc ↑
ECE ↓
AUROC ↑
Brier ↓
Acc ↑
ECE ↓
AUROC ↑
Brier ↓
CREDO (full)
.753
.094
.933
.088
.614
.052
.869
.117
Digit-exp. emission
.676
.251
.671
.218
.587
.180
.724
.229
Readout + reward only
.695
.262
.869
.252
.588
.187
.818
.186
GRPO + auxiliary head
.702
.106
.902
.111
.599
.125
.832
.155
Appendix
Table 8: Emission and training-path decomposition (seed 43). Bold marks the best in each column.
Mathematics
Code
Eval
Acc ↑
ECE ↓
AUROC ↑
Brier ↓
Acc ↑
ECE ↓
AUROC ↑
Brier ↓
RLCR
verbalized
.682
.167
.837
.181
.539
.070
.820
.144
RLCR + weighting (parsed)
verbalized
.701
.142
.837
.152
.545
.104
.818
.155
RLCR + weighting (digit-exp)
verbalized
.697
.141
.847
.148
.542
.085
.810
.149
RLCR
digit-exp
.682
.164
.855
.180
.539
.078
.822
.144
RLCR + weighting (parsed)
digit-exp
.701
.151
.857
.152
.545
.108
.823
.153
Appendix
Table 9: Verbalized baselines with discrepancy weighting (seed 43). The Eval column indicates the confidence source at evaluation. Shaded rows evaluate through the digit-probability expectation; CREDO through its logit readout; unshaded rows through the parsed verbalized numeral. Bold marks the best in each column.
Mathematics
Code
rans
Acc ↑
ECE ↓
AUROC ↑
Brier ↓
Acc ↑
ECE ↓
AUROC ↑
Brier ↓
βF+a
.753
.094
.933
.088
.614
.052
.869
.117
F(β+a)
.714
.105
.910
.110
.589
.038
.871
.113
a
.723
.113
.924
.091
.583
.042
.889
.102
Appendix
Table 10: Answer-reward shapes on seed 43; all else follows Table B.2 .
Figure 6: Calibration curves, pooled over the three seeds and over each domain’s evaluation sets. Baselines report through their verbalized numeral, CREDO through its readout. Equal-mass bins; bars are Wilson 95% intervals ( Wilson, 1927 ) .
DeepScaleR
MATH-500
AIME-24
AIME-25
AIME-26
AMC-23
AMC-24
Average
Accuracy ↑
Base
.603 ± .005
.821 ± .004
.240 ± .013
.172 ± .009
.136 ± .002
.599 ± .010
.503 ± .008
.439 ± .001
GRPO
.836 ± .006
.939 ± .001
.610 ± .021
.451 ± .021
.475 ± .030
.856 ± .014
.779 ± .014
.706 ± .009
RLCR
.821 ± .011
.932 ± .006
.564 ± .021
.382 ± .010
.465 ± .039
.851 ± .021
.762 ± .008
.682 ± .010
DCPO
.845 ± .003
.944 ± .007
.624 ± .035
.478 ± .009
.521 ± .019
.864 ± .018
.780 ± .029
.722 ± .008
Base-d
.603 ± .005
.821 ± .004
.240 ± .013
.172 ± .009
.136 ± .002
.599 ± .010
.503 ± .008
.439 ± .001
Appendix
Table 11: Per-set results on mathematics.
Accuracy ↑
ECE ↓
DeepCoder
HumanEval+
LCB v5
LCB v6
Average
DeepCoder
HumanEval+
LCB v5
LCB v6
Average
GRPO
.389 ± .015
.852 ± .003
.400 ± .002
.338 ± .007
.495 ± .004
.418 ± .058
.086 ± .011
.360 ± .042
.396 ± .057
.315 ± .041
RLCR
.388 ± .006
.858 ± .003
.400 ± .009
.365 ± .007
.502 ± .002
.311 ± .027
.054 ± .017
.271 ± .017
.276 ± .022
.228 ± .014
DCPO
.394 ± .022
.861 ± .004
.415 ± .025
.363 ± .019
.508 ± .014
.436 ± .059
.153 ± .017
.344 ± .047
.387 ± .049
.330 ± .040
GRPO-d
.389 ± .015
.852 ± .003
.400 ± .002
.338 ± .007
.495 ± .004
.405 ± .040
.050 ± .002
.345 ± .031
.391 ± .042
.298 ± .028
RLCR-d
.388 ± .006
.858 ± .003
.400 ± .009
.365 ± .007
.502 ± .002
.315 ± .022
.056 ± .011
.270 ± .016
.285 ± .023
.232 ± .013
Appendix
Table 12: Per-set results on code, models trained on mathematics.
Accuracy ↑
ECE ↓
DeepCoder
HumanEval+
LCB v5
LCB v6
Average
DeepCoder
HumanEval+
LCB v5
LCB v6
Average
Base
.318 ± .004
.829 ± .007
.345 ± .012
.330 ± .007
.455 ± .001
.558 ± .005
.105 ± .008
.503 ± .012
.506 ± .001
.418 ± .001
GRPO
.475 ± .006
.882 ± .011
.522 ± .008
.443 ± .007
.581 ± .007
.229 ± .070
.113 ± .098
.187 ± .083
.199 ± .036
.182 ± .070
RLCR
.449 ± .008
.872 ± .019
.480 ± .017
.412 ± .016
.553 ± .013
.049 ± .014
.111 ± .017
.068 ± .010
.071 ± .016
.075 ± .005
DCPO
.477 ± .016
.878 ± .016
.525 ± .013
.422 ± .010
.576 ± .003
.167 ± .031
.112 ± .015
.178 ± .041
.156 ± .013
.153 ± .024
Base-d
.318 ± .004
.829 ± .007
.345 ± .012
.330 ± .007
.455 ± .001
.528 ± .004
.074 ± .017
.474 ± .012
.481 ± .003
.389 ± .002
Appendix
Table 13: Per-set results on code.
DeepScaleR
MATH-500
AIME-24
AIME-25
AIME-26
AMC-23
AMC-24
Average
Accuracy ↑
GRPO
.743 ± .004
.892 ± .005
.374 ± .016
.276 ± .009
.288 ± .004
.728 ± .003
.644 ± .010
.564 ± .003
RLCR
.701 ± .016
.873 ± .009
.324 ± .009
.233 ± .026
.207 ± .016
.688 ± .024
.579 ± .025
.515 ± .011
DCPO
.721 ± .009
.894 ± .002
.357 ± .010
.240 ± .024
.253 ± .013
.727 ± .004
.644 ± .017
.548 ± .003
GRPO-d
.743 ± .004
.892 ± .005
.374 ± .016
.276 ± .009
.288 ± .004
.728 ± .003
.644 ± .010
.564 ± .003
RLCR-d
.701 ± .016
.873 ± .009
.324 ± .009
.233 ± .026
.207 ± .016
.688 ± .024
.579 ± .025
.515 ± .011
Appendix
Table 14: Per-set results on mathematics, models trained on code.
Mathematics
Code
Acc ↑
ECE ↓
AUROC ↑
Brier ↓
Acc ↑
ECE ↓
AUROC ↑
Brier ↓
Weighting strength κ
0
.721
.122
.904
.102
.572
.051
.843
.132
0.25
.753
.094
.933
.088
.584
.044
.854
.124
0.5
.749
.093
.924
.101
.614
.052
.869
.117
1
.730
.108
.924
.100
.591
.049
.858
.124
Appendix
Table 15: Every evaluated setting of the three coefficients of § 3 , all on seed 43.
Scaling test-time computation with reinforcement learning (RL) has emerged as a reliable path to improve large language models (LLM) reasoning ability. Yet, outcome-based reward often incentivizes models to be overconfident, leading to hallucinations, unreliable confidence-based control, and unnecessary compute allocation. We introduce Reinforcement Learning with Confidence Margin (\textbf{RLCM}), a calibration-aware RL framework that jointly optimizes correctness and confidence reliability via a margin-enhanced process reward over intermediate-budget completions. Rather than aligning confidence to correctness likelihoods, RLCM encourages to widen the confidence margin between correct and incorrect steps within a single reasoning trajectory. Across mathematical, code, logic and science benchmarks, our method substantially improves calibration while maintaining or improving accuracy. We further show that, with calibrated confidence signals, the resulting models enable more efficient conformal risk control and effective confidence-weighted aggregation.
Reinforcement learning with binary correctness rewards trains correctness, not calibrated confidence. The confidence that reasoning models verbalize is systematically overconfident, and the problem is not merely one of scale: verbalized confidence tracks how willing a model is to commit to an answer, not how likely the answer is to be right. Post-hoc rescaling therefore fits one distribution but rarely transfers. We look inside the model instead. On factual question answering, a linear probe on the hidden state between the chain of thought and the answer is substantially better calibrated: its expected calibration error is 5 to 38 times lower than that of the verbalized score across four benchmarks and two model families. However, when used to pick among N sampled answers, that same probe nearly ties majority voting yet falls far short of the oracle. Internal states answer "how certain am I" well and "which answer is right" poorly, so the signal should be reported as a confidence rather than used to select answers. As a result, we introduce probe-guided self-distillation (Probe-SD): score a model's own sampled traces with the probe, overwrite the confidence each trace states, and finetune the base checkpoint of the same family, so nothing but the model itself remains at test time. On Qwen3-14B, Probe-SD cuts ECE from 0.178 to 0.024 in-domain and from 0.542 to 0.113 out-of-domain, where it also beats post-hoc recalibration and self-consistency distillation. The resulting confidence is well-calibrated and useful for weighted voting, behaviors previously attributed to online RL, here obtained with supervised finetuning alone.
Yadong Xi, Rongsheng Zhang, Tangjie Lv +2
NetEase · Amazon · Singapore University of Technology and Design
Decision models such as Jev answer questions with probabilities, which are only useful if they are calibrated. Open-source reproductions rely on supervised fine-tuning plus temperature scaling, while reinforcement learning from verifiable rewards (RLVR) makes reasoning models overconfident. We present a working implementation of reinforcement learning for calibrated decisions (RLCD) for reasoning models: the model samples a rationale, and we score the answer distribution it commits to afterwards with a strictly proper scoring rule. A variance identity shows that scoring the mixture of several samples rewards disagreeing rationales, and that RLVR is exactly this mixture objective without its diversity term. Optimized naively, the per-rationale objective either switches reasoning off or is drowned out by policy-gradient noise, which leads to a two-stage recipe: calibrate, then reinforce. With Qwen3-1.7B on two reasoning tasks (3 seeds, paired tests), RLCD matches or beats SFT, RFT/STaR and GRPO (each temperature-scaled) in accuracy and beats all of them in selective prediction; on GSM8K answer verification a single query decides \gvTwoCovFive% of the items at ≤5% error, versus \gvGrpoCovFive% for GRPO. When uncertainty comes from annotator disagreement, RLCD provably cannot beat cross-entropy. Code and results: https://github.com/ZimmyGao/openjev-rlcd.