While Large Reasoning Models (LRMs) excel at complex reasoning, alignment through reinforcement learning often induces systemic overconfidence. In production environments, where logits may be unavailable, robust black-box uncertainty quantification (UQ) is essential for trustworthiness and safety. Focusing on question-answering for LRMs, we show that existing black-box methods, such as paraphrase-based self-consistency and confidence verbalization, offer little to no improvement over simple repeated sampling, suggesting that alignment suppresses useful output variability. We introduce prompt-level relaxation operators that broaden the model's effective output distribution by approximating the effect of an optimal policy obtained with a stronger KL-regularization parameter, hence closer to the reference model. Theoretically, we demonstrate that relaxation improves calibration. We propose Jailbreak for Uncertainty (J4U), a jailbreak-derived technique for UQ that empirically reproduces the behavioral signatures predicted by our relaxation theory. Across 3 datasets and 4 LRMs, including a closed-source production model, J4U's improvement over repeated sampling achieves statistical significance in up to 6 times more LRM-dataset-metric settings than the strongest black-box UQ state-of-the-art baseline we evaluate, with average ECE reductions up to 5 times larger. These results provide a practical tool for UQ in black-box LRM deployment.
Figures & tables
VC
Rephrase
J4U-Suffix
J4U-Prog
J4U-Art
( Xiong et al., 2024 )
( Yang et al., 2024a )
(ours)
(ours)
(ours)
Acc. ↑
-1.2% -0.005 Δ
+1.4% -0.025 Δ
+2.8% -0.001 Δ
-1.3% -0.017 Δ
+7.0% -0.008 Δ
ECE ↓
+8.6% +0.035 Δ
-3.1% -0.018 Δ
-2.9% -0.008 Δ
-15.8% -0.068 Δ
-13.8% -0.066 Δ
Brier ↓
+3.5% +0.008 Δ
+2.0% +0.002 Δ
-1.7% -0.006 Δ
-4.4% -0.015 Δ
-4.8% -0.018 Δ
NLL ↓
+5.0% +0.55 Δ
-2.0% -0.24 Δ
-6.4% -0.30 Δ
-14.4% -1.73 Δ
-14.5% -1.80 Δ
Table 1: Comparison versus RS (across 12 LRM-dataset pairs). Left: relative (%) and absolute ( Δ ) average differences; bold values denote the best-performing approach per metric. Variation per LRMs and dataset is encoded as a 4×3 heatmap (rows = LRMs, columns = datasets): green = improvement, red = degradation; darker shades indicate statistical significance.
LRM
J4U-Suffix
J4U-Prog
J4U-Art
Qwen3-4B
+7.6% +0.015 Δ
+102.9% +0.308 Δ
+72.5% +0.210 Δ
gpt-oss-20b
+9.2% +0.030 Δ
+39.9% +0.155 Δ
+17.4% +0.069 Δ
DeepSeek-R1-32B
+13.2% +0.035 Δ
+57.7% +0.158 Δ
+38.6% +0.106 Δ
GPT-5.6 Luna
+17.0% +0.018 Δ
+57.9% +0.109 Δ
+62.4% +0.114 Δ
Global Mean
+11.8% +0.024 Δ
+64.6% +0.182 Δ
+47.7% +0.125 Δ
Table 2: Increase in entropy vs πRL , in relative change (in %) and absolute difference ( Δ ).
LRM
J4U-Suffix
J4U-Prog
J4U-Art
Qwen3-4B
-5.7% -0.037 Δ
-16.8% -0.113 Δ
-28.8% -0.187 Δ
gpt-oss-20b
-1.1% -0.013 Δ
-22.3% -0.183 Δ
-31.5% -0.250 Δ
DeepSeek-R1-32B
-3.7% -0.017 Δ
-4.8% -0.037 Δ
-23.7% -0.147 Δ
Global Mean
-3.5% -0.022 Δ
-14.6% -0.111 Δ
-28.0% -0.195 Δ
Table 3: Reduction in reasoning traces cosine similarity vs πRL , in relative change (in %) and absolute difference ( Δ ), with the exclusion of GPT-5.6 Luna (reasoning traces are not accessible).
Appendix figures & tables9 assets
Supplementary material from the paper’s appendix.
Appendix
Dataset
Method
Accuracy ↑
ECE ↓
Brier ↓
NLL ↓
SynthPAI
RS
0.56 ± 0.03
0.34 ± 0.03
0.37 ± 0.03
8.27 ± 0.88
VC
0.57 ± 0.03
0.33 ± 0.03
0.36 ± 0.03
8.30 ± 0.85
Rephrase
0.56 ± 0.03
0.32 ± 0.03
0.36 ± 0.02
7.30 ± 0.88
J4U-Suffix
0.56 ± 0.03
0.30 ± 0.03
0.35 ± 0.02
6.19 ± 0.83
J4U-Prog
0.58 ± 0.03
0.23 ± 0.03
0.31 ± 0.02
3.91 ± 0.64
J4U-Art
0.57 ± 0.03
0.25 ± 0.03
0.32 ± 0.02
4.22 ± 0.64
Appendix
Table 4: Comparison of UQ techniques on Qwen3-4B . Bold indicates statistically significant improvements w.r.t. RS by bootstrap test with α=0.05 .
Dataset
Method
Accuracy ↑
ECE ↓
Brier ↓
NLL ↓
SynthPAI
RS
0.61 ± 0.03
0.26 ± 0.03
0.31 ± 0.02
5.41 ± 0.74
VC
0.59 ± 0.03
0.28 ± 0.03
0.32 ± 0.03
5.85 ± 0.78
Rephrase
0.57 ± 0.03
0.26 ± 0.03
0.32 ± 0.02
4.72 ± 0.71
J4U-Suffix
0.60 ± 0.03
0.22 ± 0.03
0.30 ± 0.02
3.63 ± 0.64
J4U-Prog
0.60 ± 0.03
0.24 ± 0.03
0.31 ± 0.02
4.75 ± 0.72
J4U-Art
0.60 ± 0.03
0.24 ± 0.03
0.30 ± 0.02
4.21 ± 0.65
Appendix
Table 5: Comparison of UQ techniques on gpt-oss-20b . Bold indicates statistically significant improvements w.r.t. RS by bootstrap test with α=0.05 .
Dataset
Method
Accuracy ↑
ECE ↓
Brier ↓
NLL ↓
SynthPAI
RS
0.57 ± 0.03
0.30 ± 0.03
0.35 ± 0.02
7.25 ± 0.83
VC
0.57 ± 0.03
0.31 ± 0.03
0.35 ± 0.02
7.30 ± 0.82
Rephrase
0.53 ± 0.03
0.32 ± 0.03
0.35 ± 0.02
5.63 ± 0.75
J4U-Suffix
0.57 ± 0.03
0.28 ± 0.03
0.33 ± 0.02
6.01 ± 0.78
J4U-Prog
0.58 ± 0.03
0.25 ± 0.03
0.32 ± 0.02
4.71 ± 0.65
J4U-Art
0.58 ± 0.03
0.25 ± 0.03
0.31 ± 0.02
4.41 ± 0.62
Appendix
Table 6: Comparison of UQ techniques on DeepSeek-R1-32B . Bold indicates statistically significant improvements w.r.t. RS by bootstrap test with α=0.05 . Red indicates significant deterioration.
Dataset
Method
Accuracy ↑
ECE ↓
Brier ↓
NLL ↓
SynthPAI
RS
0.61 ± 0.03
0.32 ± 0.03
0.35 ± 0.03
9.86 ± 0.99
VC
0.59 ± 0.03
0.37 ± 0.04
0.39 ± 0.03
11.67 ± 1.2
Rephrase
0.57 ± 0.03
0.36 ± 0.03
0.38 ± 0.03
10.61 ± 0.96
J4U-Suffix
0.60 ± 0.03
0.30 ± 0.03
0.33 ± 0.03
7.94 ± 0.98
J4U-Prog
0.59 ± 0.03
0.30 ± 0.03
0.34 ± 0.02
8.63 ± 0.94
J4U-Art
0.61 ± 0.03
0.26 ± 0.03
0.32 ± 0.03
7.45 ± 0.84
Appendix
Table 7: Comparison of UQ techniques on GPT-5.6 Luna . Bold indicates statistically significant improvements w.r.t. RS by bootstrap test with α=0.05 . Red indicates significant deterioration.
LRM
Method
Accuracy ↑
ECE ↓
NLL ↓
Qwen3-4B
RS
0.02 ± 0.01
0.72 ± 0.03
8.04 ± 1.16
VC
0.05 ± 0.02
0.73 ± 0.03
8.17 ± 1.22
Rephrase
0.02 ± 0.01
0.61 ± 0.02
8.17 ± 0.97
J4U-Suffix
0.03 ± 0.01
0.68 ± 0.03
8.32 ± 1.22
J4U-Prog
0.02 ± 0.01
0.62 ± 0.02
8.60 ± 0.98
J4U-Art
0.01 ± 0.01
0.53 ± 0.02
8.36 ± 0.99
Appendix
Table 8: Accuracy and confidence calibration on HLE (exact match). Bold indicates statistically significant improvements w.r.t. RS by bootstrap test with α=0.05 . Red indicates significant deterioration.
Qwen3-4B
gpt-oss-20b
DeepSeek-R1-32B
GPT-5.6 Luna
Dataset
Method
Mean
Var.
p-value
Mean
Var.
p-value
Mean
Var.
p-value
Mean
Var.
p-value
SynthPAI
RS
0.22
0.07
0.30
0.08
0.26
0.08
0.14
0.06
Rephrase
0.26
0.08
7 ×10−7
0.35
0.07
7 ×10−11
0.31
0.08
5 ×10−12
0.15
0.07
6 ×10−2
J4U-Suffix
0.28
0.07
3 ×10−19
0.37
0.07
1 ×10−17
0.30
0.08
2 ×10−11
0.22
0.08
7 ×10−17
J4U-Prog
0.39
0.07
2 ×10−35
0.34
0.08
1 ×10−5
0.33
0.07
1 ×10−10
0.21
0.08
1 ×10−11
J4U-Art
0.38
0.07
3 ×10−32
0.35
0.07
4 ×10−8
0.35
0.07
2 ×10−17
0.22
0.08
8 ×10−16
Appendix
Table 9: Mean and variance of Shannon entropy across policies, together with p-value of one-side Mann–Whitney U test against RS.
Qwen3-4B
gpt-oss-20b
DeepSeek-R1-32B
Dataset
Method
Mean
Var.
Mean
Var.
Mean
Var.
SynthPAI
RS
0.67
0.15
0.85
0.09
0.73
0.14
Rephrase
0.63
0.16
0.87
0.09
0.72
0.15
J4U-Suffix
0.53
0.15
0.77
0.13
0.65
0.15
J4U-Prog
0.48
0.14
0.65
0.15
0.61
0.16
J4U-Art
0.34
0.07
0.45
0.12
0.35
0.14
Appendix
Table 10: Mean and variance of average cosine similarity between CoT embeddings across policies.
Methods
Complexity
Tokens
Wall-clock
In (%)
Out (%)
Total (%)
Time (%)
VC
O(2K)
+530
+64
+135
+66.29
Rephrase
O(2K)
+141
+56
+63
+25.55
J4U-SUFFIX
O(K)
+6
+0
+0
-7.23
J4U-PROG
O(K)
+10
+17
+11
-8.14
J4U-ART
O(K)
+81
+21
+27
+5.37
Appendix
Table 11: Relative overhead compared to Repeated Sampling in O(K) .
LRM
Method
Accuracy ↑
ECE ↓
NLL ↓
Qwen3-4B
RS
0.96 ± 0.01
0.04 ± 0.01
0.71 ± 0.31
VC
0.96 ± 0.01
0.04 ± 0.01
0.65 ± 0.29
Rephrase
0.95 ± 0.01
0.04 ± 0.01
0.61 ± 0.31
NOISE 10%
0.94 ± 0.02
0.18 ± 0.02
0.66 ± 0.29
NOISE 20%
0.69 ± 0.04
0.19 ± 0.03
0.91 ± 0.27
J4U-Suffix
0.96 ± 0.01
0.04 ± 0.01
0.64 ± 0.28
Appendix
Table 12: Accuracy and confidence calibration on gsm8k where RS exhibit high accuracy and confidence calibration. Red indicates statistically significant deterioration.
Current answering paradigms for Large Reasoning Models (LRMs) often fail to account for the fact that some questions may lie beyond the model's operational capability boundary, leading to long but unproductive reasoning. In this paper, we study whether LRMs expose early signals predictive of such cases, and whether these signals can be used to mitigate unproductive reasoning. In black-box settings, we find that reasoning expressions contain failure-predictive signals. In white-box settings, we show that the hidden states of the last input token contain information that is predictive of whether a question will not be solved correctly under our evaluation setup. Building on these observations, we propose two test-time monitoring strategies: reasoning expression monitoring and hidden states monitoring, that reduce token usage by 62.7-93.6%, substantially improving efficiency and reliability while largely preserving accuracy.
As the chain-of-thought reasoning capabilities of large language models improve, evaluating and calibrating their reasoning confidence is becoming increasingly important for quantifying the uncertainty of their answers. Current methods for estimating the confidence of large language models are generally based on probabilities of selected key tokens, but the underlying mechanism remains unclear. Our pilot study finds that replacing selected token probabilities with coarse substitutes can also improve calibration, motivating us to further explore effective signals of model confidence. We introduce Divergent Token Confidence (DTC), a framework that estimates confidence by counting tokens at which two models strongly disagree during decoding. DTC identifies these divergent tokens using the Jensen-Shannon divergence between next-token distributions evaluated along the same reasoning trajectory. We find that their count is almost negatively associated with answer accuracy, thereby serving as a simple yet effective signal for uncertainty quantification. DTC supports both white-box and black-box evaluation using auxiliary models, without explicit training and affecting the generation process. Experiments across multiple model families and six mathematical benchmarks demonstrate improved calibration over probability-based and verbalized baselines. Under white-box evaluation, the count-only estimator achieves an average expected calibration error of 13.0%, compared with 32.7%-42.4% for standard full-sequence confidence methods. In black-box settings, it also improves calibration over the original verbalized scores. For example, mean expected calibration error falls from 32.1%-40.2% to 13.7%-16.3% on DeepSeek-V3.2. These findings provide new insights for improving reasoning uncertainty quantification in large language models. The code is released at https://github.com/szu-tera/DTC.git.
Feiyang Li, Shengjing Liu, Qi Zhan +7
College of Computer Science and Software Engineering, Shenzhen University · Tsinghua University · Behavioral and Spatial AI Lab, Peking University & Tongji University
Reliable uncertainty communication is critical to the trustworthiness of LLMs, yet faithful calibration (FC)--the alignment between models' intrinsic and (linguistically) expressed confidence--is a persistent failure mode. This challenge is key for large reasoning models (LRMs), whose extended reasoning traces are often interpreted by users as evidence of deliberation, competence, and confidence. Despite the importance of FC and wide usage of LRMs, the extent to which LRMs can faithfully express their confidence remains poorly understood. Moreover, the prevailing paradigm to measure FC does not generalize well to the long chain-of-thought outputs generated by LRMs, which tend to lack clear step boundaries, involve inconsistent step structure, and encode complex conditional dependencies throughout the trace--complicating estimation of intrinsic confidence. To address this challenge, we introduce a novel framework to systematically quantify FC of LRMs. Our framework analyzes linguistic decisiveness relative to three sources of internal uncertainty, based on token probabilities, hidden states, and sampled response consistency. We also devise a prefix-conditioned sampling approach to control for conditional and structural variation across traces. Applying our framework to a diverse suite of leading models, datasets, and prompts, we find that faithful confidence expression is a significant challenge for LRMs. Reasoning behaviors do not automatically translate to improved FC, and prompt interventions for non-reasoning models do not improve faithfulness in the reasoning setting. Different confidence estimators further produce divergent assessments of the same traces, revealing fragility in prior evaluation methodologies. Taken together, our work establishes FC as a distinct reliability and alignment target for LRMs, particularly as such systems are increasingly deployed in high-stakes contexts.
Areeb Gani, Asal Meskin, Gabrielle Kaili-May Liu +1