Recent work on hallucination detection in large language models has shown that, for a fixed pre-trained model and reasoning task, it is possible to estimate the model's confidence in the correctness of its outputs. Such uncertainty estimates have primarily been used to improve truthfulness by detecting or filtering confabulations. In this work, we ask whether these signals can instead be used more proactively to directly improve the accuracy of model-generated answers. We propose USteer, a simple, training-free steering mechanism that adjusts a model's layer-wise activations during inference using the gradient of a confidence measure with respect to the activations. This procedure nudges generation toward outputs with lower uncertainty at inference time, without modifying model parameters or requiring additional supervision. We show that this approach consistently reduces hallucination across a range of tasks, demonstrating that confidence signals can be leveraged not only for detection, but also for effective inference-time control of model behavior.
Figures & tables
Figure 1 : Overview of USteer, a training-free uncertainty-guided steering method. (a) Scheme Illustration. We intervene on output of intermediate transformer layers, with the steering mechanism highlighted in yellow. USteer is training-free, where steering directions are derived from uncertainty signals defined by the language modeling head. Under the logit lens, we evaluate uncertainty scores based on NCI ( Liu et al., 2026 ) ) and steer representations toward regions associated with lower uncertainty. (b) Example Outputs. USteer reduces hallucination, as illustrated on an example from the GSM8K dataset using Llama-3.2-1B-Instruct .
Figure 2 : Uncertainty Signals Emerge at Intermediate Layers of LLM. For each layer, we compute the score from Equation 2 across steps under the logit lens and compare its distribution for non-hallucinated (correct) and hallucinated (incorrect) examples. Correct intermediate embeddings tends to exhibit higher confidence guidance score. While the distinction between correct and incorrect is strongest at the final layer (Layer 15), intermediate layers (Layer 14 and Layer 13) already show such trend. This demonstrates that the language head provides meaningful uncertainty signals before the final layer, motivating steering at intermediate representations. Evaluation is conducted on GSM8K using Llama-3.2-1B-Instruct .
Methods
Training-free
Latency ↓
CSQA
StrategyQA
GSM8K
AQuA
Model: Llama-3.2-1B-Instruct
Standard
-
12.5
50.94
58.52
34.19
37.80
DoLa
✓
17.2
51.19
60.70
36.39
31.50
ITI
✗
22.1
54.46
59.39
35.56
37.80
TruthX
✗
21.9
53.07
61.57
34.87
36.22
USteer (Ours)
✓
12.8
54.05
63.32
36.24
39.37
Table 1 : USteer mitigates hallucination across datasets and models with negligible latency overhead. Performance is reported in reasoning task accuracy and latency (ms/token; lower is better). The best performance is shown in bold , and the second-best performance is underlined .
Confidence Region
USteer
DoLa
ITI
TruthX
All examples
+2.6
+0.7
+2.1
+1.3
Bottom 25%
+6.7
+3.8
+5.1
+3.3
Bottom 5%
+14.4
+8.5
+9.8
+9.8
Table 2 : USteer provides larger accuracy gains in low-confidence regions. Confidence regions are pooled across CSQA, StrategyQA, GSM8K, and AQuA. Net accuracy gains are reported in percentage points over standard greedy decoding on Llama-3.2-1B-Instruct .
Figure 5
Standard
Lower Layers
Middle Layers
Upper Layers
Accuracy (%)
68.56
69.00
69.43
69.43
Table 3 : USteer consistently mitigates hallucination across different selections of steering layers. Experiments on StrategyQA using Qwen-2.5-3B-Instruct . Accuracy reported. USteer improves performance when applied at lower, middle, and upper layers.
temp = 0.2
temp = 0.5
temp = 0.8
temp = 1.0
Standard
50.38±0.46
50.22±0.29
47.83±0.32
45.44±0.38
USteer (ours)
53.55±0.38
52.22±0.42
50.37±0.19
48.78±0.32
Table 4 : USteer mitigates hallucination under stochastic decoding , consistently yielding gains across a range of temperatures with statistical significance. Results are reported on CSQA using Llama-3.2-1B-Instruct , with accuracy shown as the mean and 95% confidence intervals.
Method
CSQA
StrategyQA
GSM8K
AQuA
Standard
77.72
68.56
79.91
54.72
USteer w/ NCI
79.20
70.74
80.97
55.12
USteer w/ fDBD
78.21
71.18
81.20
55.12
USteer w/ softmax
77.89
69.43
80.52
55.12
Table 5 : USteer improves decoding performance with alternative confidence guidance scores. We report accuracy (%) on Qwen-2.5-3B-Instruct . Both fDBD and softmax confidence improve performance over standard inference across all datasets.
Appendix figures & tables8 assets
Supplementary material from the paper’s appendix.
Appendix
Methods
CSQA
StrategyQA
GSM8K
AQuA
Standard
84.44
72.49
91.51
75.20
DoLa
84.03
73.36
92.19
75.98
ITI
84.19
75.98
91.58
73.80
TruthX
84.44
74.24
91.81
74.24
USteer (Ours)
85.26
75.55
91.74
77.17
Appendix
Table 6 : USteer consistantly mitigates hallucination and improves reasoning accuracy on Qwen-2.5-14B-Instruct . Performance is reported in reasoning task accuracy. The best performance is shown in bold , and the second-best performance is underlined .
Methods
CSQA
StrategyQA
GSM8K
AQuA
Standard
86.49
78.60
94.54
77.17
USteer
86.90
79.04
95.00
77.17
Appendix
Table 7 : USteer consistently mitigates hallucination and improves reasoning accuracy on Qwen-3-32B . We report task accuracy. USteer consistently improves standard decoding across datasets.
Method
CSQA
StrategyQA
GSM8K
AQuA
( n=1221 )
( n=229 )
( n=1319 )
( n=254 )
Model: Llama-3.2-1B-Instruct
Standard
50.94±2.78
58.52±6.55
34.19±2.58
37.80±5.91
DoLa
51.19±2.83
60.70±6.33
36.39±2.58
31.50±5.51
ITI
54.46±2.87
59.39±6.33
35.56±2.54
37.80±5.91
TruthX
53.07±2.78
61.57±6.33
34.87±2.58
36.22±5.91
Appendix
Table 8 : Accuracy (%) with 95% bootstrap confidence intervals under greedy decoding. The confidence intervals quantify sampling uncertainty arising from the finite evaluation sets. Their widths are largely determined by dataset size, with wider intervals observed on the smaller StrategyQA and AQuA evaluation sets.
Llama-3.2-1B-Instruct
Qwen-2.5-3B-Instruct
Qwen-2.5-14B-Instruct
L15 / L14 / L13
L35 / L34 / L33
L47 / L46 / L45
GSM8K
74.9 / 69.2 / 65.3
75.4 / 75.3 / 65.9
71.2 / 74.3 / 62.8
CSQA
60.8 / 56.5 / 55.3
62.8 / 58.9 / 57.6
69.1 / 68.7 / 65.5
StrategyQA
52.3 / 50.8 / 51.6
61.3 / 55.2 / 57.1
74.6 / 73.7 / 67.7
AQuA
50.3 / 52.9 / 53.8
76.5 / 69.6 / 62.3
75.6 / 77.9 / 62.9
Appendix
Table 9 : Intermediate-layer uncertainty guidance scores are predictive of generation correctness across models, datasets, and layers. The layer-wise uncertainty guidance score g(zl) from Definition 3.1 is used as the prediction score, with correct generations treated as the positive class. Each entry reports AUROC for the final three layers, ordered from the final layer to the two preceding layers. A higher AUROC indicates stronger discrimination between correct and incorrect generations.
Temperature
0.2
0.5
0.8
1.0
Standard
96.3%
90.6%
83.7%
77.6%
USteer
97.1%
92.4%
86.4%
81.1%
Appendix
Table 10 : USteer modestly increases the selection of the most likely token. We report the percentage of generation steps at which the token with the highest predictive probability is selected under different sampling temperatures.
Temperature
0.2
0.5
0.8
1.0
Standard
0.767
0.751
0.718
0.675
USteer
0.800
0.786
0.756
0.719
Appendix
Table 11 : USteer produces more confident generation trajectories. We report the average predictive probability of the sampled tokens under different sampling temperatures.
Temperature
0.2
0.5
0.8
1.0
Standard
1.52
1.86
2.15
2.37
USteer
1.36
1.68
1.96
2.20
Appendix
Table 12 : USteer produces more consistent answers across repeated sampling runs. We report the average number of distinct answers generated across five runs for each question. Lower values indicate less variation among the generated answers.
Temperature
0.0
0.2
0.5
0.8
1.0
Standard
34.81
34.89
35.25
36.01
37.16
USteer
35.89
35.81
35.96
36.42
37.83
Appendix
Table 13 : Generation length exhibits a similar temperature-dependent trend with and without USteer. We report the average generation length under greedy decoding ( T=0 ) and stochastic decoding at different sampling temperatures.
We propose a lightweight and single-pass uncertainty quantification method for detecting hallucinations in Large Language Models. The method uses attention matrices to estimate uncertainty without requiring repeated sampling or external models. Specifically, we measure the Kullback-Leibler divergence between each attention head's distribution and a uniform reference distribution, and use these features in a logistic regression probe. Across multiple datasets, task types, and model families, attention divergence is highly predictive of answer correctness and performs competitively with existing uncertainty estimation methods. We find that this signal is concentrated in middle layers and on factual tokens such as named entities and numbers, suggesting that attention dynamics provides an efficient and interpretable white-box signal of model uncertainty.
Large language models (LLMs) are prone to hallucinations, i.e., statements unsupported by the input or training data, hindering reliable deployment. In parallel, numerous uncertainty estimation (UE) methods have been proposed to quantify model confidence and are often implicitly treated as proxies for model failure. However, the relationship between uncertainty and hallucinations remains insufficiently characterized. We present a systematic empirical study of the association between uncertainty estimators and hallucinations in LLMs. Rather than assuming this association, we evaluate directly when and to what extent it holds. We consider a diverse set of uncertainty estimators, including information-theoretic, sampling-based, and reflexive estimators, and examine their behavior across hallucination settings. Our experiments cover both intrinsic hallucinations (violations of input faithfulness) and extrinsic hallucinations (unsupported claims relative to training data), using four complementary benchmarks, including RAGTruth and HalluLens. We find that the association is highly variable and often weak, depending on the hallucination type and the LLM under evaluation. These results challenge the use of uncertainty as a direct signal of hallucination and clarify when it provides actionable information.
Yedidia Agnimo, Anna Korba, Annabelle Blangero +2
Ekimetrics · Centre Inria de l’Université Grenoble Alpes · CREST, ENSAE +1
We introduce CAROL (Chain-based Adaptive Reconfiguration Over Lattices), a probabilistic framework for test-time hallucination reduction in large language models. Rather than relying on token-level uncertainty, CAROL defines a semantic uncertainty measure based on the consistency between generated responses and a trusted context, inducing a string-submodular objective over a lattice of textual sequences. This formulation enables hallucination mitigation to be cast as a Markov chain accept-reject process with provable convergence and near-optimality guarantees, allowing the model to iteratively refine outputs toward semantic consistency. By operating at the level of meaning, CAROL unifies hallucination detection and mitigation within a single framework. Empirical results on question answering and multi-agent reasoning benchmarks show that CAROL significantly reduces hallucinations and improves reliability and interpretability compared to likelihood-based and retrieval-augmented baselines, while maintaining competitive computational efficiency.
Joan Vendrell Gallart, Solmaz Kia, Russell Bent +1
Department of Mechanical and Aerospace University of California Irvine Irvine, CA 92617-4322, USA · T-5 Los Alamos National Laboratory Los Alamos, NM 88220, USA · CAI-4 Los Alamos National Laboratory Los Alamos, NM 88220, USA