Organizations: College of Computer Science and Software Engineering, Shenzhen University · Tsinghua University · Behavioral and Spatial AI Lab, Peking University & Tongji University
As the chain-of-thought reasoning capabilities of large language models improve, evaluating and calibrating their reasoning confidence is becoming increasingly important for quantifying the uncertainty of their answers. Current methods for estimating the confidence of large language models are generally based on probabilities of selected key tokens, but the underlying mechanism remains unclear. Our pilot study finds that replacing selected token probabilities with coarse substitutes can also improve calibration, motivating us to further explore effective signals of model confidence. We introduce Divergent Token Confidence (DTC), a framework that estimates confidence by counting tokens at which two models strongly disagree during decoding. DTC identifies these divergent tokens using the Jensen-Shannon divergence between next-token distributions evaluated along the same reasoning trajectory. We find that their count is almost negatively associated with answer accuracy, thereby serving as a simple yet effective signal for uncertainty quantification. DTC supports both white-box and black-box evaluation using auxiliary models, without explicit training and affecting the generation process. Experiments across multiple model families and six mathematical benchmarks demonstrate improved calibration over probability-based and verbalized baselines. Under white-box evaluation, the count-only estimator achieves an average expected calibration error of 13.0%, compared with 32.7%-42.4% for standard full-sequence confidence methods. In black-box settings, it also improves calibration over the original verbalized scores. For example, mean expected calibration error falls from 32.1%-40.2% to 13.7%-16.3% on DeepSeek-V3.2. These findings provide new insights for improving reasoning uncertainty quantification in large language models. The code is released at https://github.com/szu-tera/DTC.git.
Figures & tables
Figure 1: Divergent-token selection and its relationship to path accuracy. (a) A generator produces a reasoning trajectory. (b) A token is selected when the JSD between two models’ next-token distributions exceeds θ . (c) Path accuracy decreases as the number of divergent tokens increases under single- and dual-auxiliary probing.
Table 1: Uncertainty quantification performance (ECE) in white-box settings. In each row, the best result is in bold and the second-best one is underlined (excluding PRM). †: UQAC variant by ours.
Figure 3: Calibration plots and probability histogram for Qwen2.5-14B on MATH-500. The x -axis shows mean confidence within 20 probability bins. The calibration curve ( blue line with μ±σ ) displays actual accuracy per bin, while the gray shadow represents the probability proportion.
Method
AIME24
AIME25
HMMT25
HMMT26
Acc ↑
ECE ↓
AUC ↑
Acc ↑
ECE ↓
AUC ↑
Acc ↑
ECE ↓
AUC ↑
Acc ↑
ECE ↓
AUC ↑
Default CoT
72.2
–
–
60.9
–
–
46.1
–
–
50.9
–
–
→ + PRM
–
39.0
88.2
–
39.5
80.4
–
41.1
60.4
–
38.1
75.1
→ + DTClin
–
11.2
74.6
–
0 9.4
78.7
–
12.7
67.9
–
12.5
72.7
→ + DTCprod
–
24.6
79.4
–
23.9
82.1
–
29.1
75.0
–
26.4
76.5
Verb. Conf.
67.4
40.8
85.5
55.3
39.6
86.8
38.3
40.3
78.2
45.0
40.1
84.3
Table 2: Uncertainty quantification performance with DeepSeek-V3.2 in black-box settings.
Figure 4: Sensitivity of DTClin to θ and auxiliary size, single-auxiliary Qwen2.5 on AMC23. We search θ in steps of 0.05 and report the θ (a) with the lowest ECE (b) and its AUROC (c).
Figure 5: DTClin sensitivity to θ with auxiliary fixed at 1.5B and generating models ∈{7,14,32} B (solid: AMC23; dashed: AIME24). The shaded band marks the white-box operating point θ=0.50 .
Method
AIME24
AIME25
HMMT25
HMMT26
ECE ↓
AUC ↑
ECE ↓
AUC ↑
ECE ↓
AUC ↑
ECE ↓
AUC ↑
Verb. Conf.
40.8
85.5
39.6
86.8
40.3
78.2
40.1
84.3
w/ DTC
0 6.3
85.8
0 3.7
87.7
10.5
77.8
0 6.2
84.2
Verb. TopK
32.0
88.4
32.6
85.7
31.6
79.3
32.0
83.6
w/ DTC
17.0
90.3
14.6
87.3
18.8
79.1
15.7
83.6
Verb. PD
32.6
86.2
33.2
84.9
31.5
74.3
32.3
83.1
Table 3: Verbalized vs. w/ DTC on DeepSeek-V3.2.
Appendix figures & tables17 assets
Supplementary material from the paper’s appendix.
Appendix
White-box
Black-box
Benchmarks
MATH-500 ( Lightman et al., 2024 ) AMC23 ( Balunovic et al., 2025 ) AIME24 ( Balunovic et al., 2025 ) AIME25 ( Balunovic et al., 2025 )
AIME24 ( Balunovic et al., 2025 ) AIME25 ( Balunovic et al., 2025 ) HMMT25 ( Balunovic et al., 2025 ) HMMT26 ( Balunovic et al., 2025 )
Generators
Qwen2.5-7B-Instruct ( Qwen Team, 2024 ) Qwen2.5-14B-Instruct ( Qwen Team, 2024 ) Qwen2.5-32B-Instruct ( Qwen Team, 2024 ) Qwen3-8B ( Yang et al., 2025 ) Qwen3-14B ( Yang et al., 2025 ) Qwen3-32B ( Yang et al., 2025 ) Gemma3-12B-IT ( Gemma Team, 2025 ) Gemma3-27B-IT ( Gemma Team, 2025 )
DeepSeek-V3.2 ( DeepSeek-AI, 2025 ) Qwen3-30B-A3B-Instruct-2507 ( Yang et al., 2025 ) Qwen3-4B-Instruct-2507 ( Yang et al., 2025 )
Auxiliaries
Qwen2.5-1.5B-Instruct ( Qwen Team, 2024 ) Qwen3-1.7B ( Yang et al., 2025 ) Gemma3-4B-IT ( Gemma Team, 2025 )
Table 5: Uncertainty quantification performance (AUROC, ↑ higher is better) in white-box settings. In each row, the best result is in bold and the second-best one is underlined (excluding PRM). †: UQAC variant by ours.
Method
AIME24
AIME25
HMMT25
HMMT26
Acc ↑
ECE ↓
AUC ↑
Acc ↑
ECE ↓
AUC ↑
Acc ↑
ECE ↓
AUC ↑
Acc ↑
ECE ↓
AUC ↑
Default CoT
61.4
–
–
45.6
–
–
30.3
–
–
34.7
–
–
→ + PRM
–
36.0
82.9
–
36.6
80.3
–
37.8
53.3
–
33.4
78.0
→ + DTClin
–
11.3
72.9
–
0 9.0
81.5
–
13.1
69.6
–
13.0
74.6
→ + DTCprod
–
16.3
75.8
–
10.7
84.8
–
15.4
78.0
–
14.3
79.0
Verb. Conf.
61.8
45.6
80.8
46.6
45.9
85.6
27.3
45.6
81.9
33.0
45.1
79.8
Appendix
Table 6: Uncertainty quantification performance with Qwen3-4B in black-box settings.
Method
AIME24
AIME25
HMMT25
HMMT26
Acc ↑
ECE ↓
AUC ↑
Acc ↑
ECE ↓
AUC ↑
Acc ↑
ECE ↓
AUC ↑
Acc ↑
ECE ↓
AUC ↑
Default CoT
73.4
–
–
59.8
–
–
42.4
–
–
43.5
–
–
→ + PRM
–
37.5
86.0
–
37.8
78.4
–
38.7
53.6
–
34.6
77.4
→ + DTClin
–
15.9
70.2
–
12.3
75.8
–
13.1
72.1
–
0 9.6
77.8
→ + DTCprod
–
18.2
75.1
–
16.6
79.6
–
19.5
78.7
–
17.1
80.5
Verb. Conf.
75.7
44.5
83.9
60.6
44.3
80.8
41.8
42.9
85.3
43.9
43.9
79.2
Appendix
Table 7: Uncertainty quantification performance with Qwen3-30B-A3B in black-box settings.
Method
AIME24
AIME25
HMMT25
HMMT26
ECE ↓
AUC ↑
ECE ↓
AUC ↑
ECE ↓
AUC ↑
ECE ↓
AUC ↑
Verb. Conf.
44.5
83.9
44.3
80.8
42.9
85.3
43.9
79.2
w/ DTC
17.8
87.0
15.7
83.8
13.0
87.9
17.4
85.9
Verb. TopK
40.2
88.2
40.6
88.7
41.2
85.4
41.6
83.0
w/ DTC
11.1
83.0
10.4
84.9
0 5.4
86.5
0 6.1
86.4
Verb. PD
33.5
75.5
34.1
69.6
31.1
77.3
34.9
72.3
Appendix
Table 8: Verbalized vs. w/ DTC on Qwen3-30B-A3B.
Method
AIME24
AIME25
HMMT25
HMMT26
ECE ↓
AUC ↑
ECE ↓
AUC ↑
ECE ↓
AUC ↑
ECE ↓
AUC ↑
Verb. Conf.
45.6
80.8
45.9
85.6
45.6
81.9
45.1
79.8
w/ DTC
30.8
82.0
29.4
87.0
31.0
81.6
31.1
81.2
Verb. TopK
47.0
68.3
47.0
75.3
46.8
74.0
45.5
77.3
w/ DTC
33.7
68.6
32.8
75.5
34.3
73.8
30.9
76.9
Verb. PD
46.9
69.9
48.1
73.7
43.9
75.5
42.9
70.9
Appendix
Table 9: Verbalized vs. w/ DTC on Qwen3-4B.
Figure 9: Black-box sensitivity of DTClin to θ and auxiliary size. We search θ in steps of 0.05 and report the θ (a) with the lowest ECE (b) and its AUROC (c). Cells without a finished probe are left blank.
Figure 10: DTClin sensitivity to θ with A′′ fixed at 1.5B (DeepSeek-V3.2 generating; solid: AIME25; dashed: HMMT25). The shaded band marks the black-box operating point θ=0.70 .
Figure 11: Divergent-token count and ratio against path accuracy, single-auxiliary Qwen2.5-7B on MATH-500. Each curve is one auxiliary at the threshold used in Figure 1 (c).
Figure 12: Divergence choice for divergent-token selection. Each column is one disagreement score (JSD, forward KL, reverse KL) for DTClin with auxiliary fixed at 1.5B and generating models ∈{7,14,32} B (solid: AMC23; dashed: AIME24). The top row shows ECE (%) and the bottom row AUROC (%) versus θ .
Figure 13: Divergence choice for black-box divergent-token selection. Each column is one disagreement score (JSD, forward KL, reverse KL) for DTClin with auxiliaries fixed at Qwen2.5-7B and Qwen2.5-1.5B and generators Qwen3-4B, Qwen3-30B-A3B, and DeepSeek-V3.2 (solid: AIME25; dashed: HMMT25). The top row shows ECE (%) and the bottom row AUROC (%) versus θ .
Figure 14: Saturation length n for DTClin . The top row shows ECE (%) and the bottom row AUROC (%) versus θ at step 0.01 , averaged over MATH-500, AMC23, AIME24, and AIME25.
Figure 15: Saturation length n for black-box DTClin . Each column is one generator with Qwen2.5 auxiliaries 7 B and 1.5 B. The top row shows ECE (%) and the bottom row AUROC (%) versus θ at step 0.01 , averaged over AIME24, AIME25, and HMMT25.
Figure 16: Offset k for DTCprod . Each column is the smallest white-box generator in its family (Qwen2.5-7B, Qwen3-8B, Gemma3-12B) with the family auxiliary. ECE (%) is shown versus θ at step 0.01 , averaged over MATH-500, AMC23, AIME24, and AIME25.
Reliable uncertainty communication is critical to the trustworthiness of LLMs, yet faithful calibration (FC)--the alignment between models' intrinsic and (linguistically) expressed confidence--is a persistent failure mode. This challenge is key for large reasoning models (LRMs), whose extended reasoning traces are often interpreted by users as evidence of deliberation, competence, and confidence. Despite the importance of FC and wide usage of LRMs, the extent to which LRMs can faithfully express their confidence remains poorly understood. Moreover, the prevailing paradigm to measure FC does not generalize well to the long chain-of-thought outputs generated by LRMs, which tend to lack clear step boundaries, involve inconsistent step structure, and encode complex conditional dependencies throughout the trace--complicating estimation of intrinsic confidence. To address this challenge, we introduce a novel framework to systematically quantify FC of LRMs. Our framework analyzes linguistic decisiveness relative to three sources of internal uncertainty, based on token probabilities, hidden states, and sampled response consistency. We also devise a prefix-conditioned sampling approach to control for conditional and structural variation across traces. Applying our framework to a diverse suite of leading models, datasets, and prompts, we find that faithful confidence expression is a significant challenge for LRMs. Reasoning behaviors do not automatically translate to improved FC, and prompt interventions for non-reasoning models do not improve faithfulness in the reasoning setting. Different confidence estimators further produce divergent assessments of the same traces, revealing fragility in prior evaluation methodologies. Taken together, our work establishes FC as a distinct reliability and alignment target for LRMs, particularly as such systems are increasingly deployed in high-stakes contexts.
Areeb Gani, Asal Meskin, Gabrielle Kaili-May Liu +1
The ability of large language models (LLMs) to express calibrated uncertainty is important for safe deployment. Chain-of-thought (CoT) reasoning is widely used to improve accuracy and reliability, but its effect on calibration is not fully understood. We show that this picture is incomplete: in some settings, increasing the reasoning budget beyond a task-specific threshold can cause models to become systematically overconfident, assigning high confidence to incorrect answers. We call this phenomenon Calibration Drift Under Reasoning (CDUR) and study it both theoretically and empirically. We define reasoning budget B and analyze conditions under which Expected Calibration Error ECE(B) follows a non-monotonic pattern: it first decreases as reasoning corrects errors, then increases as longer reasoning produces internally consistent but incorrect explanations. We propose a Hypothesis Lock-In model based on autoregressive generation to explain this behavior. We evaluate Llama-3.1-8B and Llama-3.3-70B on 47 reasoning-trap questions across four reasoning budgets and three seeds (1,368 API calls; 574 valid responses). The 8B model shows non-monotonic calibration behavior, while results for the 70B model are limited to baseline evaluation and are inconclusive for budget-dependent effects. We introduce CABStop, a calibration-aware stopping rule that halts reasoning when confidence diverges from an auxiliary accuracy estimate. These results suggest that increasing reasoning depth does not always improve reliability and should be monitored carefully.
Prakul Sunil Hiremath, Harshit R. Hiremath
Department of Computer Science and Engineering Visvesvaraya Technological University, Belagavi · Department of Computer Science and Business System SG Balekundri Institute of Technology, Belagavi
Reasoning language models can solve increasingly complex tasks, but struggle to produce the calibrated confidence estimates necessary for reliable deployment. Existing calibration methods usually depend on labels or repeated sampling at inference time, making them impractical in many settings. We introduce a method for unsupervised confidence calibration of reasoning LLMs when only a single generation is available at inference time. Our approach uses offline sampling on unlabeled data to derive a self-consistency-based proxy target, then distills this signal into a lightweight deployment-time confidence predictor. In a broad evaluation across 5 math and question-answering tasks using 9 reasoning models, our method substantially outperforms baselines, including under distribution shift, and improves downstream performance in selective prediction and simulated downstream decision-making.