RL-ARC: Calibrating Large Reasoning Models via Reasoning-guided Uncertainty
Authors: Gukhyeon Lee, SangKeun Lee
Organizations: Department of Artificial Intelligence, Korea University, Seoul, Republic of Korea · Department of Computer Science and Engineering, Korea University, Seoul, Republic of Korea
Language models (LMs) are commonly trained with Reinforcement Learning with Verifiable Rewards (RLVR) to enhance their reasoning capabilities. However, since RLVR does not explicitly account for calibration during training, it can lead to severe calibration degradation, including overconfidence. Recent calibration-aware training methods for LMs, which incorporate objectives for uncertainty estimation into training, improve calibration but still exhibit overconfidence under distribution shift, while sacrificing reasoning performance. To this end, we propose RL-ARC, a calibration-aware training framework that jointly leverages reasoning confidence and answer confidence. Specifically, RL-ARC leverages reasoning confidence as an auxiliary signal for calibrating answer confidence, applying it as reasoning-guided regularization for correct cases and as an overconfidence penalty for incorrect cases. Comprehensive results across ID and OOD settings show that, beyond improving calibration, RL-ARC enables reasoning models to adaptively estimate confidence based on the given question without substantially sacrificing reasoning performance, thereby highlighting the importance of reasoning confidence for training reliable reasoning models.
Figures & tables
Figure 1: Comparison of reasoning training frameworks: (a) Standard reasoning training often assigns high confidence to incorrect answers (i.e., overconfidence). (b) Previous calibration-aware training reduces overconfidence to some extent, but still tends to assign high confidence to incorrect answers on OOD benchmarks. (c) Our RL-ARC consistently improves calibration while preserving accuracy gains under OOD settings.
Figure 2: Overview of RL-ARC . RL-ARC leverages reasoning confidence differently depending on whether the prediction is correct. For correct cases, RL-ARC uses reasoning confidence to regularize answer confidence, encouraging high confidence only when both the prediction and the reasoning process are correct. For incorrect cases, RL-ARC penalizes overconfidence by leveraging confidence from both the reasoning process and the prediction.
Distribution
Method
Qwen2.5 (7B)
Qwen3 (8B)
Acc. ( ↑ )
AUROC ( ↑ )
Brier ( ↓ )
ECE ( ↓ )
Acc. ( ↑ )
AUROC ( ↑ )
Brier ( ↓ )
ECE ( ↓ )
ID (Avg.)
Base
41.19
0.58
0.54
0.54
55.63
0.51
0.42
0.42
RLVR
w/ Confidence
56.86
0.45
0.41
0.40
66.58
0.66
0.31
0.32
w/ Probability
56.86
0.60
0.41
0.41
66.58
0.76
0.27
0.29
w/ Post-hoc
56.86
0.72
0.24
0.21
66.58
0.83
0.16
0.20
Table 1: Evaluation results for accuracy (%) and calibration metrics on Math (ID) and OOD benchmarks. The best and second-best results are highlighted in boldface and underlined .
Backbone
Method
ID (Avg.)
OOD (Avg.)
Acc. ( ↑ )
AUROC ( ↑ )
Brier ( ↓ )
ECE ( ↓ )
Acc. ( ↑ )
AUROC ( ↑ )
Brier ( ↓ )
ECE ( ↓ )
Qwen2.5-7B
Base
41.19
0.58
0.54
0.54
43.52
0.53
0.49
0.49
RLVR
56.86
0.45
0.41
0.40
47.58
0.50
0.50
0.50
RLCR
54.82
0.58
0.25
0.22
46.02
0.54
0.28
0.20
RL-ARC
56.53
0.60
0.23
0.18
47.27
0.54
0.26
0.16
Llama-8B (DeepSeek-R1 Distilled)
Base
33.24
0.61
0.62
0.63
45.96
0.61
0.43
0.44
Table 2: Evaluation results across different backbone models on Math (ID) and OOD benchmarks. The best and second-best results within each model are highlighted in boldface and underlined , respectively.
Training Distribution
Method
ID (Avg.)
OOD (Avg.)
Acc. ( ↑ )
AUROC ( ↑ )
Brier ( ↓ )
ECE ( ↓ )
Acc. ( ↑ )
AUROC ( ↑ )
Brier ( ↓ )
ECE ( ↓ )
Arithmetic Reasoning
Base
41.19
0.58
0.54
0.54
43.52
0.53
0.49
0.49
RLVR
56.86
0.45
0.41
0.40
47.58
0.50
0.50
0.50
RLCR
54.82
0.58
0.25
0.22
46.02
0.54
0.28
0.20
RL-ARC
56.53
0.60
0.23
0.18
47.27
0.54
0.26
0.16
Complex Reasoning
Base
39.10
0.54
0.53
0.54
45.41
0.53
0.48
0.49
Table 3: Evaluation results across different training distributions. The best and second-best results within each training distribution are highlighted in boldface and underlined , respectively.
Method
Complex Reasoning
Factual QA
Avg.
Acc. ( ↑ )
AUROC ( ↑ )
Brier ( ↓ )
ECE ( ↓ )
Acc. ( ↑ )
AUROC ( ↑ )
Brier ( ↓ )
ECE ( ↓ )
Acc. ( ↑ )
AUROC ( ↑ )
Brier ( ↓ )
ECE ( ↓ )
Base
47.95
0.55
0.44
0.43
39.10
0.51
0.55
0.56
43.52
0.53
0.49
0.49
RLVR
54.86
0.50
0.45
0.45
40.30
0.50
0.56
0.56
47.58
0.50
0.50
0.50
RLCR
51.83
0.53
0.27
0.16
40.22
0.54
0.29
0.23
46.02
0.54
0.28
0.20
RL-ARC
53.36
0.54
0.25
0.14
41.19
0.54
0.27
0.19
47.27
0.54
0.26
0.16
Table 4: Evaluation results for accuracy (%) and calibration metrics on six OOD benchmarks, including three complex reasoning and three factual question answering benchmarks. The best and second-best results are highlighted in boldface and underlined . Here, we use Qwen2.5-7B as base model.
Figure 3: Performance (solid lines) and confidence frequency (bars) across confidence bins on OOD benchmarks for Qwen2.5-7B trained with each method.
Figure 4: Risk-coverage curves across representative benchmarks. Lower selective risk indicates better confidence estimation and selective prediction performance (i.e., better confidence-aware prediction), where more accurate predictions are assigned higher confidence.
λneg
λpos
OOD (Avg.)
Acc. ( ↑ )
AUROC ( ↑ )
AURC ( ↓ )
Brier ( ↓ )
ECE ( ↓ )
✓
✗
47.11
0.53
0.46
0.25
0.15
✓
✓
47.27
0.54
0.45
0.26
0.16
✗
✓
47.51
0.55
0.45
0.28
0.21
✗
✗
46.02
0.54
0.47
0.28
0.20
Table 5: Ablation results on OOD benchmarks. λneg and λpos indicate whether the overconfidence penalty and reasoning-guided regularization are used, respectively. The best and second-best results are highlighted in boldface and underlined .
Figure 5: Comparison between performance and calibration across different auxiliary signal scales. For RL-ARC , the left and right values in parentheses denote λpos and λneg , respectively. Here, we use Qwen2.5-7B as base model.
Figure 6: Comparison of calibration improvements and shallow reasoning reduction over the base model. The y-axis reports the relative change from the base model. Here, we use Qwen2.5-7B as base model.
Appendix figures & tables20 assets
Supplementary material from the paper’s appendix.
Appendix
Hyperparameter
Value
Maximum input length
1024
Maximum output length
4096
Rollout group size
8
Rollout temperature
0.7
Epoch
1
Batch size
1536
Appendix
Table 6: Hyperparameter settings for training.
Figure 7: Performance (solid lines) and confidence frequency (bars) across confidence bins for Qwen2.5-7B trained with each method on ID (top) and OOD (bottom) benchmarks.
Figure 8: Performance (solid lines) and confidence frequency (bars) across confidence bins for Qwen3-8B trained with each method on ID (top) and OOD (bottom) benchmarks.
Method
ID (Avg.)
OOD (Avg.)
Brier ( ↓ )
ECE ( ↓ )
Brier ( ↓ )
ECE ( ↓ )
RLVR
0.41
0.40
0.50
0.50
RLCR
0.25
0.22
0.28
0.20
RL-ARC (shuffled)
0.25
0.19
0.27
0.19
RL-ARC
0.23
0.18
0.26
0.16
Appendix
Table 7: Control experiment results on the calibration gains of RL-ARC across ID and OOD benchmarks.
Method
ID (Avg.)
OOD (Avg.)
Acc. ( ↑ )
AURC ( ↓ )
ECE ( ↓ )
Acc. ( ↑ )
AURC ( ↓ )
ECE ( ↓ )
RLVR
66.58
0.22
0.32
47.24
0.44
0.48
RLCR
65.98
0.20
0.16
50.99
0.35
0.22
RL-ARC
66.10
0.16
0.15
51.61
0.34
0.17
Appendix
Table 8: Comparison of accuracy (%), AURC, and ECE under Math (ID) and OOD benchmarks.
Method
ID (Avg.)
OOD (Avg.)
Acc. ( ↑ )
AUROC ( ↑ )
Brier ( ↓ )
ECE ( ↓ )
Acc. ( ↑ )
AUROC ( ↑ )
Brier ( ↓ )
ECE ( ↓ )
RLVR
56.86
0.45
0.41
0.40
47.58
0.50
0.50
0.50
RL-ARC
w/ Arithmetic mean
53.42
0.58
0.27
0.24
48.61
0.54
0.31
0.25
w/ Geometric mean
54.04
0.58
0.26
0.25
47.65
0.53
0.30
0.23
w/ Ours
56.53
0.60
0.23
0.18
47.27
0.54
0.26
0.16
Appendix
Table 9: Comparison of strategies for integrating reasoning confidence into calibration-aware training on Math (ID) and OOD benchmarks.
Setting
ID (Avg.)
OOD (Avg.)
λneg
λpos
Acc. ( ↑ )
AUROC ( ↑ )
Brier ( ↓ )
ECE ( ↓ )
Acc. ( ↑ )
AUROC ( ↑ )
Brier ( ↓ )
ECE ( ↓ )
0.00
0.00
54.82
0.58
0.25
0.22
46.02
0.54
0.28
0.20
0.08
0.20
55.92
0.61
0.24
0.20
46.41
0.55
0.28
0.19
0.10
0.30
54.11
0.64
0.24
0.21
46.85
0.54
0.28
0.20
0.15
0.40
56.53
0.60
0.23
0.18
47.27
0.54
0.26
0.16
0.20
0.60
57.64
0.65
0.23
0.19
47.08
0.56
0.27
0.17
Appendix
Table 10: Detailed evaluation results of auxiliary signal scaling for RL-ARC across ID and OOD benchmarks. The best results are highlighted in boldface .
Figure 9: The designed prompt used to estimate verbalized confidence during reasoning.
Benchmark
Method
Acc. ( ↑ )
AUROC ( ↑ )
Brier ( ↓ )
ECE ( ↓ )
Big-Math
Base
0.50
0.55
0.46
0.46
RLVR
w/ Confidence
0.67
0.51
0.32
0.32
w/ Probability
0.67
0.61
0.31
0.31
w/ Post-hoc
0.67
0.73
0.21
0.08
Behavioral Calibration
0.67
0.50
0.22
0.03
Appendix
Table 11: Evaluation results for accuracy and calibration metrics on Math (ID) benchmarks. The best results are highlighted in boldface . Here, we use Qwen2.5-7B as base model.
Benchmark
Method
Acc. ( ↑ )
AUROC ( ↑ )
Brier ( ↓ )
ECE ( ↓ )
Big-Math
Base
0.73
0.52
0.24
0.23
RLVR
w/ Confidence
0.78
0.67
0.21
0.21
w/ Probability
0.78
0.71
0.22
0.22
w/ Post-hoc
0.78
0.79
0.16
0.13
Behavioral Calibration
0.78
0.67
0.15
0.08
Appendix
Table 12: Evaluation results for accuracy and calibration metrics on Big-Math, GSM8K, and MATH-500 (ID) benchmarks. The best results are highlighted in boldface . We use Qwen3-8B as the base model.
Benchmark
Method
Acc. ( ↑ )
AUROC ( ↑ )
Brier ( ↓ )
ECE ( ↓ )
AMC23
Base
0.68
0.40
0.32
0.32
RLVR
w/ Confidence
0.90
0.57
0.10
0.09
w/ Probability
0.90
0.79
0.11
0.10
w/ Post-hoc
0.90
0.93
0.08
0.11
Behavioral Calibration
0.78
0.76
0.15
0.10
Appendix
Table 13: Evaluation results for accuracy and calibration metrics on AMC23, AIME24, and AIME25 (ID) benchmarks. The best results are highlighted in boldface . We use Qwen3-8B as the base model.
Benchmark
Method
Acc. ( ↑ )
AUROC ( ↑ )
Brier ( ↓ )
ECE ( ↓ )
Big-Math
Base
0.29
0.61
0.65
0.67
RLVR
0.61
0.53
0.35
0.34
RLCR
0.58
0.76
0.22
0.14
RL-ARC (ours)
0.63
0.77
0.20
0.08
GSM8K
Base
0.38
0.55
0.60
0.60
RLVR
0.82
0.51
0.16
0.13
Appendix
Table 14: Evaluation results for accuracy and calibration metrics on Math (ID) benchmarks. The best results are highlighted in boldface . Here, we use Llama-8B distilled from DeepSeek-R1 as the base model.
Benchmark
Method
Acc. ( ↑ )
AUROC ( ↑ )
Brier ( ↓ )
ECE ( ↓ )
StrategyQA
Base
0.68
0.58
0.27
0.23
RLVR
w/ Confidence
0.70
0.48
0.29
0.29
w/ Probability
0.70
0.56
0.30
0.29
w/ Post-hoc
0.70
0.45
0.25
0.19
Behavioral Calibration
0.72
0.53
0.20
0.04
Appendix
Table 15: Evaluation results for accuracy and calibration metrics on complex reasoning (OOD) benchmarks. The best results are highlighted in boldface . Here, we use Qwen2.5-7B as base model.
Benchmark
Method
Acc. ( ↑ )
AUROC ( ↑ )
Brier ( ↓ )
ECE ( ↓ )
SimpleQA
Base
0.13
0.50
0.78
0.81
RLVR
w/ Confidence
0.14
0.51
0.81
0.83
w/ Probability
0.14
0.46
0.79
0.81
w/ Post-hoc
0.14
0.47
0.50
0.60
Behavioral Calibration
0.12
0.50
0.44
0.58
Appendix
Table 16: Evaluation results for accuracy and calibration metrics on factual question answering (OOD) benchmarks. The best results are highlighted in boldface . Here, we use Qwen2.5-7B as base model.
Benchmark
Method
Acc. ( ↑ )
AUROC ( ↑ )
Brier ( ↓ )
ECE ( ↓ )
GPQA
Base
0.37
0.54
0.51
0.52
RLVR
0.39
0.50
0.61
0.61
RLCR
0.40
0.54
0.29
0.22
RL-ARC (ours)
0.40
0.52
0.26
0.15
GSM8K
Base
0.73
0.52
0.25
0.22
RLVR
0.49
0.50
0.51
0.51
Appendix
Table 17: Evaluation results for accuracy and calibration metrics on complex reasoning (OOD) benchmarks. The best results are highlighted in boldface . Here, we use Qwen2.5-7B as the base model.
Benchmark
Method
Acc. ( ↑ )
AUROC ( ↑ )
Brier ( ↓ )
ECE ( ↓ )
SimpleQA
Base
0.13
0.50
0.78
0.81
RLVR
0.11
0.50
0.89
0.89
RLCR
0.12
0.54
0.33
0.45
RL-ARC (ours)
0.12
0.51
0.24
0.35
TriviaQA
Base
0.57
0.52
0.39
0.38
RLVR
0.60
0.50
0.40
0.40
Appendix
Table 18: Evaluation results for accuracy and calibration metrics on factual knowledge (OOD) benchmarks. The best results are highlighted in boldface . Here, we use Qwen2.5-7B as the base model.
Benchmark
Method
Acc. ( ↑ )
AUROC ( ↑ )
Brier ( ↓ )
ECE ( ↓ )
StrategyQA
Base
0.71
0.62
0.24
0.20
RLVR
w/ Confidence
0.70
0.56
0.27
0.26
w/ Probability
0.70
0.49
0.30
0.30
w/ Post-hoc
0.70
0.55
0.30
0.31
Behavioral Calibration
0.69
0.60
0.24
0.17
Appendix
Table 19: Evaluation results for accuracy and calibration metrics on complex reasoning (OOD) benchmarks. The best results are highlighted in boldface . Here, we use Qwen3-8B as base model.
Benchmark
Method
Acc. ( ↑ )
AUROC ( ↑ )
Brier ( ↓ )
ECE ( ↓ )
SimpleQA
Base
0.12
0.63
0.51
0.62
RLVR
w/ Confidence
0.10
0.53
0.78
0.83
w/ Probability
0.10
0.38
0.80
0.81
w/ Post-hoc
0.10
0.56
0.29
0.38
Behavioral Calibration
0.11
0.60
0.50
0.61
Appendix
Table 20: Evaluation results for accuracy and calibration metrics on factual question answering (OOD) benchmarks. The best results are highlighted in boldface . Here, we use Qwen3-8B as base model.
Benchmark
Method
Acc. ( ↑ )
AUROC ( ↑ )
Brier ( ↓ )
ECE ( ↓ )
StrategyQA
Base
0.64
0.61
0.27
0.23
RLVR
0.64
0.50
0.32
0.29
RLCR
0.64
0.61
0.24
0.14
RL-ARC (ours)
0.65
0.59
0.24
0.12
HotpotQA
Base
0.47
0.55
0.49
0.50
RLVR
0.27
0.50
0.66
0.68
Appendix
Table 21: Evaluation results for accuracy and calibration metrics on complex reasoning (OOD) benchmarks. The best results are highlighted in boldface . Here, we use Llama-8B distilled from DeepSeek-R1 as the base model.
Benchmark
Method
Acc. ( ↑ )
AUROC ( ↑ )
Brier ( ↓ )
ECE ( ↓ )
SimpleQA
Base
0.10
0.57
0.65
0.75
RLVR
0.10
0.51
0.80
0.84
RLCR
0.12
0.62
0.37
0.52
RL-ARC (ours)
0.10
0.65
0.17
0.28
NQ-Open
Base
0.40
0.68
0.46
0.49
RLVR
0.41
0.51
0.53
0.54
Appendix
Table 22: Evaluation results for accuracy and calibration metrics on factual knowledge (OOD) benchmarks. The best results are highlighted in boldface . Here, we use Llama-8B distilled from DeepSeek-R1 as the base model.
Scaling test-time computation with reinforcement learning (RL) has emerged as a reliable path to improve large language models (LLM) reasoning ability. Yet, outcome-based reward often incentivizes models to be overconfident, leading to hallucinations, unreliable confidence-based control, and unnecessary compute allocation. We introduce Reinforcement Learning with Confidence Margin (\textbf{RLCM}), a calibration-aware RL framework that jointly optimizes correctness and confidence reliability via a margin-enhanced process reward over intermediate-budget completions. Rather than aligning confidence to correctness likelihoods, RLCM encourages to widen the confidence margin between correct and incorrect steps within a single reasoning trajectory. Across mathematical, code, logic and science benchmarks, our method substantially improves calibration while maintaining or improving accuracy. We further show that, with calibrated confidence signals, the resulting models enable more efficient conformal risk control and effective confidence-weighted aggregation.
Reasoning language models are increasingly asked not only to answer difficult questions, but also to estimate their likelihood of success. Existing methods typically elicit confidence only once: either before thinking or after answering. We argue that confidence in reasoning models is state-dependent: before thinking, confidence should estimate the chance of the model correctly solving the prompt, while after thinking it should predict whether the realized answer is likely to be correct. This distinction determines the appropriate supervision target: prompt-level success should supervise confidence estimates made after seeing the prompt, while individual answer-level correctness should supervise confidence estimates made after answering. We introduce CALIBER (Calibration Before and After Reasoning), which elicits both estimates and supervises each with the target matched to its information state. Under this unified protocol, CALIBER reduces Expected Calibration Error (ECE) by 52.5% over the strongest single-confidence baseline on BigMathDigits for the 7B model, while achieving the best Brier score and AUROC, and remains within 2.1 points of the best accuracy. Further, on a larger 30B model, CALIBER achieves the best ECE on BigMathDigits while remaining competitive in Brier score and AUROC. Out of distribution, it achieves the best ECE and Brier score on GPQA and TriviaQA, and remains competitive on SimpleQA. Ablations further show that this position-target alignment is most beneficial under distribution shift where it consistently reduces calibration error across all out-of-distribution benchmarks.
Reinforcement learning with verifiable rewards (RLVR) trains reasoning models to produce correct answers, but does not ensure that their stated confidence is calibrated. The resulting models are systematically overconfident. Recent methods train calibration inside the RLVR loop by having the model state a numerical confidence alongside its answer, but they all obtain the confidence by sampling it as text. This choice imposes two costs: a sampled confidence introduces variance and in practice collapses to a handful of distinct values, and sampling makes the confidence non-differentiable, forcing the calibration loss through a scalar reward. We propose CREDO (Confidence REaDOut) to replace sampling with a deterministic readout. While RLVR optimizes correctness, CREDO reads the confidence from a dedicated token pair in the model's output distribution and trains it by differentiable regression. CREDO further turns the trained confidence into a signal for accuracy, weighting rollouts by how far confidence and outcome disagree, so that accuracy and calibration improve together. Across mathematical and code reasoning, CREDO attains the best accuracy and calibration, and the gains extend to abstention and selective prediction.
Chenxiao Fan, Chongming Gao, Gangyi Zhang +8
University of Science and Technology of China · Qwen Business Unit of Alibaba · National University of Singapore