While Large Reasoning Models (LRMs) excel at complex reasoning, alignment through reinforcement learning often induces systemic overconfidence. In production environments, where logits may be unavailable, robust black-box uncertainty quantification (UQ) is essential for trustworthiness and safety. Focusing on question-answering for LRMs, we show that existing black-box methods, such as paraphrase-based self-consistency and confidence verbalization, offer little to no improvement over simple repeated sampling, suggesting that alignment suppresses useful output variability. We introduce prompt-level relaxation operators that broaden the model's effective output distribution by approximating the effect of an optimal policy obtained with a stronger KL-regularization parameter, hence closer to the reference model. Theoretically, we demonstrate that relaxation improves calibration. We propose Jailbreak for Uncertainty (J4U), a jailbreak-derived technique for UQ that empirically reproduces the behavioral signatures predicted by our relaxation theory. Across 3 datasets and 4 LRMs, including a closed-source production model, J4U's improvement over repeated sampling achieves statistical significance in up to 6 times more LRM-dataset-metric settings than the strongest black-box UQ state-of-the-art baseline we evaluate, with average ECE reductions up to 5 times larger. These results provide a practical tool for UQ in black-box LRM deployment.
Figures & tables
VC
Rephrase
J4U-Suffix
J4U-Prog
J4U-Art
( Xiong et al., 2024 )
( Yang et al., 2024a )
(ours)
(ours)
(ours)
Acc. ↑
-1.2% -0.005 Δ
+1.4% -0.025 Δ
+2.8% -0.001 Δ
-1.3% -0.017 Δ
+7.0% -0.008 Δ
ECE ↓
+8.6% +0.035 Δ
-3.1% -0.018 Δ
-2.9% -0.008 Δ
-15.8% -0.068 Δ
-13.8% -0.066 Δ
Brier ↓
+3.5% +0.008 Δ
+2.0% +0.002 Δ
-1.7% -0.006 Δ
-4.4% -0.015 Δ
-4.8% -0.018 Δ
NLL ↓
+5.0% +0.55 Δ
-2.0% -0.24 Δ
-6.4% -0.30 Δ
-14.4% -1.73 Δ
-14.5% -1.80 Δ
Table 1: Comparison versus RS (across 12 LRM-dataset pairs). Left: relative (%) and absolute ( Δ ) average differences; bold values denote the best-performing approach per metric. Variation per LRMs and dataset is encoded as a 4×3 heatmap (rows = LRMs, columns = datasets): green = improvement, red = degradation; darker shades indicate statistical significance.
LRM
J4U-Suffix
J4U-Prog
J4U-Art
Qwen3-4B
+7.6% +0.015 Δ
+102.9% +0.308 Δ
+72.5% +0.210 Δ
gpt-oss-20b
+9.2% +0.030 Δ
+39.9% +0.155 Δ
+17.4% +0.069 Δ
DeepSeek-R1-32B
+13.2% +0.035 Δ
+57.7% +0.158 Δ
+38.6% +0.106 Δ
GPT-5.6 Luna
+17.0% +0.018 Δ
+57.9% +0.109 Δ
+62.4% +0.114 Δ
Global Mean
+11.8% +0.024 Δ
+64.6% +0.182 Δ
+47.7% +0.125 Δ
Table 2: Increase in entropy vs πRL , in relative change (in %) and absolute difference ( Δ ).
LRM
J4U-Suffix
J4U-Prog
J4U-Art
Qwen3-4B
-5.7% -0.037 Δ
-16.8% -0.113 Δ
-28.8% -0.187 Δ
gpt-oss-20b
-1.1% -0.013 Δ
-22.3% -0.183 Δ
-31.5% -0.250 Δ
DeepSeek-R1-32B
-3.7% -0.017 Δ
-4.8% -0.037 Δ
-23.7% -0.147 Δ
Global Mean
-3.5% -0.022 Δ
-14.6% -0.111 Δ
-28.0% -0.195 Δ
Table 3: Reduction in reasoning traces cosine similarity vs πRL , in relative change (in %) and absolute difference ( Δ ), with the exclusion of GPT-5.6 Luna (reasoning traces are not accessible).
Appendix figures & tables9 assets
Supplementary material from the paper’s appendix.
Appendix
Dataset
Method
Accuracy ↑
ECE ↓
Brier ↓
NLL ↓
SynthPAI
RS
0.56 ± 0.03
0.34 ± 0.03
0.37 ± 0.03
8.27 ± 0.88
VC
0.57 ± 0.03
0.33 ± 0.03
0.36 ± 0.03
8.30 ± 0.85
Rephrase
0.56 ± 0.03
0.32 ± 0.03
0.36 ± 0.02
7.30 ± 0.88
J4U-Suffix
0.56 ± 0.03
0.30 ± 0.03
0.35 ± 0.02
6.19 ± 0.83
J4U-Prog
0.58 ± 0.03
0.23 ± 0.03
0.31 ± 0.02
3.91 ± 0.64
J4U-Art
0.57 ± 0.03
0.25 ± 0.03
0.32 ± 0.02
4.22 ± 0.64
Appendix
Table 4: Comparison of UQ techniques on Qwen3-4B . Bold indicates statistically significant improvements w.r.t. RS by bootstrap test with α=0.05 .
Dataset
Method
Accuracy ↑
ECE ↓
Brier ↓
NLL ↓
SynthPAI
RS
0.61 ± 0.03
0.26 ± 0.03
0.31 ± 0.02
5.41 ± 0.74
VC
0.59 ± 0.03
0.28 ± 0.03
0.32 ± 0.03
5.85 ± 0.78
Rephrase
0.57 ± 0.03
0.26 ± 0.03
0.32 ± 0.02
4.72 ± 0.71
J4U-Suffix
0.60 ± 0.03
0.22 ± 0.03
0.30 ± 0.02
3.63 ± 0.64
J4U-Prog
0.60 ± 0.03
0.24 ± 0.03
0.31 ± 0.02
4.75 ± 0.72
J4U-Art
0.60 ± 0.03
0.24 ± 0.03
0.30 ± 0.02
4.21 ± 0.65
Appendix
Table 5: Comparison of UQ techniques on gpt-oss-20b . Bold indicates statistically significant improvements w.r.t. RS by bootstrap test with α=0.05 .
Dataset
Method
Accuracy ↑
ECE ↓
Brier ↓
NLL ↓
SynthPAI
RS
0.57 ± 0.03
0.30 ± 0.03
0.35 ± 0.02
7.25 ± 0.83
VC
0.57 ± 0.03
0.31 ± 0.03
0.35 ± 0.02
7.30 ± 0.82
Rephrase
0.53 ± 0.03
0.32 ± 0.03
0.35 ± 0.02
5.63 ± 0.75
J4U-Suffix
0.57 ± 0.03
0.28 ± 0.03
0.33 ± 0.02
6.01 ± 0.78
J4U-Prog
0.58 ± 0.03
0.25 ± 0.03
0.32 ± 0.02
4.71 ± 0.65
J4U-Art
0.58 ± 0.03
0.25 ± 0.03
0.31 ± 0.02
4.41 ± 0.62
Appendix
Table 6: Comparison of UQ techniques on DeepSeek-R1-32B . Bold indicates statistically significant improvements w.r.t. RS by bootstrap test with α=0.05 . Red indicates significant deterioration.
Dataset
Method
Accuracy ↑
ECE ↓
Brier ↓
NLL ↓
SynthPAI
RS
0.61 ± 0.03
0.32 ± 0.03
0.35 ± 0.03
9.86 ± 0.99
VC
0.59 ± 0.03
0.37 ± 0.04
0.39 ± 0.03
11.67 ± 1.2
Rephrase
0.57 ± 0.03
0.36 ± 0.03
0.38 ± 0.03
10.61 ± 0.96
J4U-Suffix
0.60 ± 0.03
0.30 ± 0.03
0.33 ± 0.03
7.94 ± 0.98
J4U-Prog
0.59 ± 0.03
0.30 ± 0.03
0.34 ± 0.02
8.63 ± 0.94
J4U-Art
0.61 ± 0.03
0.26 ± 0.03
0.32 ± 0.03
7.45 ± 0.84
Appendix
Table 7: Comparison of UQ techniques on GPT-5.6 Luna . Bold indicates statistically significant improvements w.r.t. RS by bootstrap test with α=0.05 . Red indicates significant deterioration.
LRM
Method
Accuracy ↑
ECE ↓
NLL ↓
Qwen3-4B
RS
0.02 ± 0.01
0.72 ± 0.03
8.04 ± 1.16
VC
0.05 ± 0.02
0.73 ± 0.03
8.17 ± 1.22
Rephrase
0.02 ± 0.01
0.61 ± 0.02
8.17 ± 0.97
J4U-Suffix
0.03 ± 0.01
0.68 ± 0.03
8.32 ± 1.22
J4U-Prog
0.02 ± 0.01
0.62 ± 0.02
8.60 ± 0.98
J4U-Art
0.01 ± 0.01
0.53 ± 0.02
8.36 ± 0.99
Appendix
Table 8: Accuracy and confidence calibration on HLE (exact match). Bold indicates statistically significant improvements w.r.t. RS by bootstrap test with α=0.05 . Red indicates significant deterioration.
Qwen3-4B
gpt-oss-20b
DeepSeek-R1-32B
GPT-5.6 Luna
Dataset
Method
Mean
Var.
p-value
Mean
Var.
p-value
Mean
Var.
p-value
Mean
Var.
p-value
SynthPAI
RS
0.22
0.07
0.30
0.08
0.26
0.08
0.14
0.06
Rephrase
0.26
0.08
7 ×10−7
0.35
0.07
7 ×10−11
0.31
0.08
5 ×10−12
0.15
0.07
6 ×10−2
J4U-Suffix
0.28
0.07
3 ×10−19
0.37
0.07
1 ×10−17
0.30
0.08
2 ×10−11
0.22
0.08
7 ×10−17
J4U-Prog
0.39
0.07
2 ×10−35
0.34
0.08
1 ×10−5
0.33
0.07
1 ×10−10
0.21
0.08
1 ×10−11
J4U-Art
0.38
0.07
3 ×10−32
0.35
0.07
4 ×10−8
0.35
0.07
2 ×10−17
0.22
0.08
8 ×10−16
Appendix
Table 9: Mean and variance of Shannon entropy across policies, together with p-value of one-side Mann–Whitney U test against RS.
Qwen3-4B
gpt-oss-20b
DeepSeek-R1-32B
Dataset
Method
Mean
Var.
Mean
Var.
Mean
Var.
SynthPAI
RS
0.67
0.15
0.85
0.09
0.73
0.14
Rephrase
0.63
0.16
0.87
0.09
0.72
0.15
J4U-Suffix
0.53
0.15
0.77
0.13
0.65
0.15
J4U-Prog
0.48
0.14
0.65
0.15
0.61
0.16
J4U-Art
0.34
0.07
0.45
0.12
0.35
0.14
Appendix
Table 10: Mean and variance of average cosine similarity between CoT embeddings across policies.
Methods
Complexity
Tokens
Wall-clock
In (%)
Out (%)
Total (%)
Time (%)
VC
O(2K)
+530
+64
+135
+66.29
Rephrase
O(2K)
+141
+56
+63
+25.55
J4U-SUFFIX
O(K)
+6
+0
+0
-7.23
J4U-PROG
O(K)
+10
+17
+11
-8.14
J4U-ART
O(K)
+81
+21
+27
+5.37
Appendix
Table 11: Relative overhead compared to Repeated Sampling in O(K) .
LRM
Method
Accuracy ↑
ECE ↓
NLL ↓
Qwen3-4B
RS
0.96 ± 0.01
0.04 ± 0.01
0.71 ± 0.31
VC
0.96 ± 0.01
0.04 ± 0.01
0.65 ± 0.29
Rephrase
0.95 ± 0.01
0.04 ± 0.01
0.61 ± 0.31
NOISE 10%
0.94 ± 0.02
0.18 ± 0.02
0.66 ± 0.29
NOISE 20%
0.69 ± 0.04
0.19 ± 0.03
0.91 ± 0.27
J4U-Suffix
0.96 ± 0.01
0.04 ± 0.01
0.64 ± 0.28
Appendix
Table 12: Accuracy and confidence calibration on gsm8k where RS exhibit high accuracy and confidence calibration. Red indicates statistically significant deterioration.
College of Computer Science and Software Engineering, Shenzhen University · Tsinghua University · Behavioral and Spatial AI Lab, Peking University & Tongji University