Organizations: State Key Laboratory of AI Safety · Institute of Computing Technology, Chinese Academy of Sciences · University of Chinese Academy of Sciences
Reliable self-assessment is essential for large language models (LLMs), yet they often remain highly confident when their answers are wrong. We study \emph{concurrent confidence calibration}, where confidence is learned alongside capability improvement rather than calibrated only after training. Reinforcement learning from verifiable rewards (RLVR) provides a natural setting for this paradigm, as it continuously produces responses paired with verifiable correctness feedback. Existing concurrent methods, however, learn both capability and confidence through reinforcement learning within shared policy parameters, potentially coupling two fundamentally different learning problems. We instead propose \emph{shared experience but separate learning}: capability and confidence learn from the same trajectories, but through separate optimization mechanisms and parameters. Based on this principle, we introduce \textbf{CoCal (Companion Confidence Calibration)}, which trains a lightweight companion from rollout hidden states and verifier-derived correctness supervision while leaving task optimization unchanged. Experiments on Qwen3-8B and Qwen3-14B show that CoCal improves confidence estimation without sacrificing task performance, outperforming both RL-based concurrent methods and matched post-hoc calibration. The learned companion further generalizes across domains and policy shifts, while the benefits of CoCal persist at both scales.
Figures & tables
Figure 1: Comparison of post-hoc calibration, coupled concurrent training, and CoCal. CoCal learns from shared RLVR experience while separating capability and confidence learning.
Figure 2: Overview of CoCal. Shared RLVR rollouts support task learning through policy optimization and confidence learning through direct supervision of a separate companion.
Method
MATH500
AIME24
AIME25
Acc. ↑
MeanConf
ECE ↓
AUROC ↑
Acc. ↑
MeanConf
ECE ↓
AUROC ↑
Acc. ↑
MeanConf
ECE ↓
AUROC ↑
Base + Verbal
79.8
0.995
0.196
0.566
25.0
0.972
0.718
0.587
18.7
0.977
0.790
0.730
Verbal
85.2
0.997
0.059
0.548
43.0
0.956
0.498
0.877
33.0
0.963
0.646
0.865
SelfConsis
86.0
0.868
0.026
0.945
47.3
0.576
0.118
0.937
34.7
0.603
0.257
0.869
TokenProb
86.0
0.819
0.087
0.834
47.3
0.668
0.246
0.873
34.7
0.638
0.330
0.943
P(True)
86.0
0.988
0.128
0.861
47.3
0.960
0.486
0.908
34.7
0.981
0.634
0.927
Table 1: Accuracy and confidence quality on five mathematical datasets using Qwen3-8B after 120 training steps. Base + Verbal uses the pre-RL model; other training-free baselines use the GRPO-trained model. Bold and underlined values indicate the best and second-best results, respectively.
Figure 3: Reliability diagrams and confidence histograms on AMC24 after 120 training steps. Upper panels show bin accuracy and the ideal-calibration diagonal; lower panels show response counts.
Figure 4: Task accuracy, confidence, calibration, and discrimination across policy-training stages on five mathematical datasets. CoCal results average five seeds; both variants share the same task accuracy.
Method
Training
Inference
LLM
Sampling
MLP
LLM
MLP
(h)
(h)
(s)
(s)
(ms)
CoCal
16.67
–
15.1
0.372
0.199
PostCal
16.67
14.22
≈15.1
0.372
0.199
SelfConsis
16.67
–
–
3.225
–
Table 2: Training and per-question inference costs. LLM inference latency is averaged over MATH500 questions using vLLM on a B200 GPU.
Method
SQuAD
WebQuestions
ComplexWebQuestions
Acc. ↑
MeanConf
ECE ↓
AUROC ↑
Acc. ↑
MeanConf
ECE ↓
AUROC ↑
Acc. ↑
MeanConf
ECE ↓
AUROC ↑
Base + Verbal
30.7
0.923
0.620
0.647
53.9
0.983
0.442
0.608
33.2
0.947
0.625
0.632
Verbal
31.1
0.907
0.603
0.656
53.7
0.978
0.441
0.623
33.0
0.918
0.596
0.677
SelfConsis
31.7
0.749
0.432
0.687
57.1
0.755
0.211
0.712
34.2
0.643
0.301
0.754
TokenProb
31.7
0.871
0.554
0.667
57.1
0.883
0.313
0.648
34.2
0.881
0.539
0.759
P(True)
31.7
0.845
0.554
0.666
57.1
0.886
0.374
0.643
34.2
0.809
0.501
0.694
Table 3: Math-to-factual transfer on HonestyBench using Qwen3-8B after 120 training steps. Base + Verbal uses the pre-RL model; other training-free baselines use the GRPO-trained model. Bold and underlined values indicate the best and second-best results, respectively.
Figure 8
Appendix figures & tables12 assets
Supplementary material from the paper’s appendix.
Appendix
Setting
Value
Policy training
Training steps
120
Questions per step
256
Responses per question
8
Maximum response length (GRPO / DCPO)
3,000
Maximum response length (RLCR)
4,096
Appendix
Table 4: Hyperparameters for policy training, calibrator training, and evaluation. Response lengths are measured in tokens.
Method
MATH500
AIME24
AIME25
Acc. ↑
MeanConf
ECE ↓
AUROC ↑
Acc. ↑
MeanConf
ECE ↓
AUROC ↑
Acc. ↑
MeanConf
ECE ↓
AUROC ↑
RLCR
85.8
0.963
0.098
0.747
37.0
0.762
0.392
0.860
25.3
0.784
0.530
0.895
DCPO
86.6
0.994
0.105
0.706
37.7
0.937
0.556
0.766
28.0
0.940
0.660
0.875
CoCal (group)
87.0
0.784
0.093
0.909
39.3
0.392
0.068
0.890
30.3
0.368
0.124
0.891
CoCal (binary)
87.0
0.831
0.055
0.910
39.3
0.478
0.126
0.887
30.3
0.446
0.166
0.902
Appendix
Table 5: Accuracy and confidence quality on five mathematical datasets using Qwen3-8B after 40 training steps. Bold and underlined values indicate the best and second-best results, respectively.
Method
MATH500
AIME24
AIME25
Acc. ↑
MeanConf
ECE ↓
AUROC ↑
Acc. ↑
MeanConf
ECE ↓
AUROC ↑
Acc. ↑
MeanConf
ECE ↓
AUROC ↑
RLCR
88.0
0.897
0.058
0.835
37.7
0.356
0.092
0.936
28.3
0.330
0.049
0.964
DCPO
87.0
0.977
0.069
0.684
41.7
0.754
0.338
0.870
33.7
0.741
0.405
0.930
CoCal (group)
89.0
0.822
0.075
0.898
44.0
0.414
0.104
0.875
36.7
0.389
0.099
0.929
CoCal (binary)
89.0
0.860
0.051
0.905
44.0
0.547
0.160
0.855
36.7
0.517
0.185
0.943
Appendix
Table 6: Accuracy and confidence quality on five mathematical datasets using Qwen3-8B after 80 training steps. Bold and underlined values indicate the best and second-best results, respectively.
Figure 7: Confidence estimation using training data produced by early steps, evaluated on the fixed step-40 GRPO policy. Points average five seeds and five mathematical datasets; error bars show seed standard deviations.
Figure 8: Reliability diagrams for Qwen3-8B after 120 policy-training steps on all five mathematical datasets. Colored bars show accuracy within nonempty confidence bins; hatched regions show the absolute gap to the actual bin-mean confidence. The dashed diagonal is an ideal-calibration reference. Lower panels report bin counts.
Figure 9: Confidence performance across matched rollout-data budgets. Both methods use binary supervision; results average five mathematical datasets.
Figure 10: Training-data composition and mean test confidence under matched raw rollout budgets. Both methods use binary supervision. Left: the fraction of incorrect effective training responses, after excluding validation and truncated responses. Right: mean confidence on the fixed 120-step test responses, averaged over five companion seeds and five mathematical datasets. The dashed line marks macro task accuracy.
Method
MATH500
AIME24
AIME25
Acc. ↑
MeanConf
ECE ↓
AUROC ↑
Acc. ↑
MeanConf
ECE ↓
AUROC ↑
Acc. ↑
MeanConf
ECE ↓
AUROC ↑
Base + Verbal
84.0
0.993
0.151
0.717
30.7
0.964
0.654
0.728
23.7
0.967
0.729
0.800
Verbal
88.8
0.991
0.103
0.747
44.3
0.949
0.506
0.712
33.7
0.961
0.624
0.820
SelfConsis
90.8
0.925
0.045
0.863
49.0
0.565
0.096
0.939
37.0
0.548
0.178
0.868
TokenProb
90.8
0.958
0.050
0.813
49.0
0.886
0.396
0.847
37.0
0.867
0.497
0.897
P(True)
90.8
0.925
0.078
0.842
49.0
0.667
0.222
0.844
37.0
0.601
0.280
0.869
Appendix
Table 7: Accuracy and confidence quality on five mathematical datasets using Qwen3-14B after 120 training steps. Base + Verbal uses the pre-RL model; other training-free baselines use the GRPO-trained model. Bold and underlined values indicate the best and second-best results, respectively.
Feature
Qwen3-8B
Qwen3-14B
ECE ↓
AUROC ↑
ECE ↓
AUROC ↑
Last token
0.093
0.900
0.087
0.898
Response mean
0.095
0.871
0.104
0.869
Last + mean
0.096
0.896
0.082
0.897
Appendix
Table 8: Confidence performance with different token aggregations, averaged over five seeds and five mathematical datasets.
Figure 11: Average confidence estimation performance using internal states from different layers on mathematical datasets.
Setting
Math
HonestyBench
ECE ↓
AUROC ↑
ECE ↓
AUROC ↑
Pre-generation
0.124
0.785
0.268
0.650
Post-generation (group)
0.103
0.884
0.104
0.725
Appendix
Table 9: Pre- and post-generation confidence performance of CoCal on Qwen3-8B after 120 policy-training steps. Results average five seeds and five datasets.
Dataset
Pre (ms)
Post (s)
Post/Pre
MATH500
38.78
58.46
1508×
AIME 2024
43.76
230.11
5258×
AIME 2025
41.78
184.56
4417×
AMC 2023
24.71
62.94
2548×
AMC 2024
25.22
79.72
3162×
Appendix
Table 10: Mean time to confidence on eight questions per dataset.
Scaling test-time computation with reinforcement learning (RL) has emerged as a reliable path to improve large language models (LLM) reasoning ability. Yet, outcome-based reward often incentivizes models to be overconfident, leading to hallucinations, unreliable confidence-based control, and unnecessary compute allocation. We introduce Reinforcement Learning with Confidence Margin (\textbf{RLCM}), a calibration-aware RL framework that jointly optimizes correctness and confidence reliability via a margin-enhanced process reward over intermediate-budget completions. Rather than aligning confidence to correctness likelihoods, RLCM encourages to widen the confidence margin between correct and incorrect steps within a single reasoning trajectory. Across mathematical, code, logic and science benchmarks, our method substantially improves calibration while maintaining or improving accuracy. We further show that, with calibrated confidence signals, the resulting models enable more efficient conformal risk control and effective confidence-weighted aggregation.
Training large language models (LLMs) with reinforcement learning (RL) has significantly advanced their performance on reasoning and question-answering tasks. However, prevailing RL reward designs typically prioritize response correctness, neglecting to incentivize models to express their confidence accurately. This leads to a critical problem: performance gains are often accompanied by poor calibration between confidence and accuracy, misleading models to overconfidently hallucinate when uncertain. To address this limitation, we propose Correctness and Confidence Calibration Reinforcement Learning (C3RL), a novel RL algorithm integrating correctness, calibration and dataset-informed reference accuracy rewards together. Comprehensive evaluation across 8 text and multimodal datasets demonstrates that C3RL enhances calibration without sacrificing accuracy, outperforming the current state-of-the-art method in both performance and calibration metrics. Utilizing the well-calibrated verbalized confidence from C3RL, we further introduce Confidence-based Adaptive Test Time Scaling (CAS), an adjustable inference-time strategy that allocates computational resources based on response confidence. Experiments show that CAS surpasses majority voting on both in-domain and out-of-domain datasets while reducing the inference budget by up to 12.33 times. We believe the synergy of C3RL and CAS paves the way for deploying more reliable and resource-efficient LLMs. The code, data and models will be released.
Xuqing Yang, Yi Yuan, Shanzhe Lei +1
Shanghai Jiao Tong University · Southeast University · Shanghai AI Laboratory
Accurate confidence estimation is critical for reliable use of large language models (LLMs). Prior work on LLM calibration largely focuses on response-level confidence, which estimates the correctness of a single generated output. However, this formulation is misaligned with many practical settings where the central question is how likely a model is to solve a query overall. We show that this mismatch results from the stochastic nature of modern LLM decoding, under which single-response correctness fails to reflect underlying model capability. To address this issue, we introduce capability calibration, a new evaluation framework for measuring how well query-level confidence aligns with a model's expected accuracy on individual queries. We formally distinguish capability calibration (CC) from response calibration (RC) and show that the two differ both theoretically and empirically. We further show that CC is better suited than RC to applications like pass@k prediction and inference budget allocation. Finally, we evaluate common confidence estimation methods to understand the practical feasibility of CC.