Large Language Models (LLMs) have achieved significant breakthroughs across various domains, but they can still produce unreliable or misleading outputs. For responsible LLM applications, uncertainty quantification techniques are used to estimate a model's uncertainty about its outputs, indicating the likelihood that those outputs may be problematic. For LLM reasoning tasks, it is essential to estimate uncertainty not only in the final answer but also in the intermediate reasoning process, particularly to identify where uncertainty arises. Such information may enable more fine-grained and targeted interventions during inference. In this study, we investigate which metrics can effectively localize uncertain places within an LLM reasoning trajectory. Our study reveals that uncertain intermediate continuations are more likely to occur at tokens that are highly sensitive to perturbations in the embeddings of preceding tokens. In our experiments, we show that such perturbation-based metrics achieve stronger performance in localizing uncertain intermediate steps than baseline methods, including probability-based, sampling-based, and Bayesian-based approaches. Meanwhile, our proposed metrics also enjoy good simplicity and efficiency.
Figures & tables
Figure 1: A response of Llama in logical reasoning ( Suzgun et al., 2023 ) . The model first states that ‘Alice is now right midfielder’ and ‘Alice and Claire trade positions’ , it then incorrectly infers ‘Claire is now left midfielder.’ While, the correct inference should be ‘Claire is now right midfielder’. Our perturbation-based metrics flag the token ‘left’ with high uncertainty score. Note: the scores are min-max normalized in the figure.
Figure 2: The model’s alternative tokens at key positions, and the regenerated continuations for the example in Figure 1 . We provide additional positions & examples in Figure 7 in Appendix E .
Math
Logic
Llama
Qwen
Mistral
Llama
Qwen
Mistral
10%
20%
10%
20%
10%
20%
10%
20%
10%
20%
10%
20%
NLL-max
0.19
0.30
0.21
0.38
0.16
0.28
0.21
0.41
0.27
0.50
0.15
0.32
Entropy-max
0.20
0.34
0.23
0.41
0.14
0.24
0.18
0.39
0.25
0.50
0.13
0.30
TokUR-max
0.27
0.48
0.26
0.45
0.28
0.44
0.27
0.42
0.29
0.47
0.25
0.45
NLL-avg
0.15
0.25
0.16
0.32
0.10
0.19
0.18
0.39
0.20
0.39
0.10
0.26
Table 1: Performance of each method in 300 responses with incorrect steps, The 95% conf. interval is ±0.057 for 300 Bernoulli trials with probability around 0.5.
GPQA
gpt-oss-20b
Qwen3.5-27B
10%
20%
10%
20%
NLL-max
0.21
0.30
0.23
0.38
Entropy-max
0.18
0.28
0.28
0.39
NLL-avg
0.13
0.31
0.13
0.28
Entropy-avg
0.15
0.28
0.15
0.30
Table 2: Performance on GPQA.
Math
Logic
Llama
Qwen
Llama
Qwen
10%
20%
10%
20%
10%
20%
10%
20%
NLL-max
0.22
0.29
0.14
0.20
0.20
0.30
0.10
0.19
Entropy-max
0.23
0.30
0.14
0.19
0.22
0.30
0.09
0.18
TokUR-max
0.23
0.31
0.14
0.23
0.23
0.33
0.13
0.17
Rand. Pert.
0.27
0.31
0.21
0.27
0.24
0.31
0.17
0.19
Table 3: The max re-generation risk of flagged uncertain steps (95% conf. interval is ±0.049 )
Figure 3: Time and memory efficiency.
FactScore
Llama
Qwen
Mistral
NLL
0.610
0.638
0.604
Entropy
0.635
0.623
0.596
TokUR
0.706
0.670
0.674
CCP
0.660
0.659
0.687
TokenSAR
0.716
0.631
0.641
Graph Uncert.
0.778
0.878
0.887
Table 5: Hallucination Detection.
Appendix figures & tables13 assets
Supplementary material from the paper’s appendix.
Appendix
Math
algebra
cnt. & prob.
number theory
precalculus
inter. algebra
prealgebra
Llama
5 / 5
4 / 5
5 / 5
4 / 5
5 / 5
5 / 5
Qwen
4 / 5
4 / 5
4 / 5
4 / 5
5 / 5
5 / 5
Mistral
4 / 5
5 / 5
4 / 5
5 / 5
5 / 5
5 / 5
Logic
logical deduction
tracking shuffled
web of lies
formal fallacies
navigate
temp. seq.
Appendix
Table 6: Human examination result for ‘first wrong step’ annotation in 180 responses.
Figure 4: Prompts used in the experiments. Title of each prompt is in bold, placeholders are marked as ‘[]’ with description within.
Math
Logic
Llama
Qwen
Mistral
Llama
Qwen
Mistral
10%
20%
10%
20%
10%
20%
10%
20%
10%
20%
10%
20%
NLL-max
0.18
0.29
0.21
0.39
0.15
0.28
0.21
0.39
0.28
0.51
0.17
0.33
Entropy-max
0.19
0.33
0.24
0.41
0.14
0.24
0.21
0.40
0.26
0.51
0.15
0.31
TokUR-max
0.27
0.49
0.26
0.44
0.28
0.44
0.27
0.42
0.29
0.45
0.26
0.45
NLL-avg
0.13
0.24
0.17
0.31
0.09
0.18
0.18
0.38
0.19
0.39
0.11
0.27
Appendix
Table 7: Performance of each method in 300 responses with incorrect final answer, The 95% conf. interval is ≈±0.057 for 300 Bernoulli trials (p=0.5).
Math
Logic
Llama
Qwen
Mistral
Llama
Qwen
Mistral
10%
20%
10%
20%
10%
20%
10%
20%
10%
20%
10%
20%
NLL-max
0.20
0.33
0.21
0.42
0.17
0.32
0.20
0.39
0.26
0.47
0.14
0.32
Entropy-max
0.24
0.37
0.24
0.40
0.12
0.24
0.18
0.38
0.23
0.47
0.15
0.31
TokUR-max
0.25
0.46
0.27
0.45
0.28
0.45
0.26
0.40
0.29
0.49
0.23
0.45
NLL-avg
0.17
0.27
0.16
0.33
0.11
0.20
0.17
0.38
0.19
0.37
0.10
0.27
Appendix
Table 8: Results on MATH and BBH, following the same setting in Section 4.1 , with LLM temperature T=0.5 .
Table 9: Experiment results using different system prompts.
Figure 5: Result under various hyper-parameters.
Figure 6: For the same setting as Figure 2 , we re-generate continuations conditioned on adversarial perturbed embeddings H^1:t−1 . We find that the continuations are also coherent and on-task, similar to Figure 2 . This shows that adversarial perturbation on the embedding also does not disrupt reasoning coherence.
Figure 7: Extra cases of regeneration from perturbed conditions.
Figure 8: Comparison of un-normalized perturbation-based scores between uncertain responses (red) and certain responses (blue). Here certain responses indicate self-consistency >0.7 .
Figure 9: Extra case studies on math reasoning task. Each uncertainty score is min-max normalized over the complete output to fit the figure, and only the critical segments are shown.
Figure 10: Extra case studies on logical reasoning task. Each uncertainty score is min-max normalized over the complete output to fit the figure, and only the critical segments are shown.
Figure 11: Case studies on logical reasoning task. Each uncertainty score is min-max normalized over the complete output to fit the figure, and only part of tokens are shown.
Figure 12: Correct but risky case. The LLM is performing a correct mathematical derivation from y=−4−72x−76 to y=−72x−734 . The token ‘34’ is flagged as uncertain though correct. Re-generation from this step reveals risk score of 0.6, where 0.5 of the risk is due to “ y=−72x−732 ” where model incorrectly generate ‘32’ instead of ‘34’ , and 0.1 is because model generate “ y=−72x−738 ”, both re-generations show the uncertainty of token ‘34’ .
Data Science Section, IT University of Copenhagen · Pioneer Centre for Artificial Intelligence · MaiNLP, Center for Information and Language Processing, LMU Munich +2