The wide adoption of LLMs across broad NLG applications heightens the importance of providing users with the means to avert errors and hallucinations. Uncertainty quantification is poised to fill that gap; with low uncertainty (high confidence), as a proxy for correctness, allowing users to be selective (e.g., reject low-confidence, likely incorrect responses). Correlation between confidence and correctness then serves as a useful criterion for evaluation of uncertainty quantifiers (UQs). But in NLG, where diverse responses can be adequate to a prompt, obtaining reliable correctness judgements is not simple, especially without human intervention. Errors in automated judgement are hardly avoidable and known to diminish the reliability of evaluation protocols (Santilli et al., 2025; Ielanskyi et al., 2025). In a meta-analysis of published work, we show that automated judgement is the present norm. Besides, automated judgements are rarely validated against human ones, and the validation of the UQ evaluation they automate is even rarer. With experiments in question answering, using 4 LLMs, human and automated judgements and 7 popular UQs, we find that i) a judge's performance can only coarsely predict the observed impact of its errors on the reliability of UQ evaluation, and that ii) judgement errors tend to misrepresent informative UQs most. We link these observations to patterns of correlation between confidence and categories of judgement error.
Figures & tables
Figure 1: An informative UQ (top-left) assigns higher confidence to responses that humans deem correct (C) than to those that humans deem wrong (W), allowing a decision maker to be selective ( e.g. , by rejecting lower-confidence responses wrt a threshold (dashed line)). If assigned confidences do not neatly separate Cs from Ws (top-right), the UQ leads to selection errors (shown in red). The degree of separation gives us a criterion to choose amongst UQs. But when judgement is automated (bottom), as routinely done in UQ evaluation, judgement errors (C → W or W → C; shown as diamonds) distort our view of the relative merits of the UQs. This affects informative UQs most: on the left, we go from 0/6 to 3/6 selection errors (2 rejected Cs and 1 accepted W), while on the right we go from 3/6 to 4/6 errors.
Correctness judge
# Papers
Proportion
LLM judge
26
38.8%
Overlap based
19
28.3%
Exact/fuzzy match
14
20.9%
Semantic similarity
11
16.4%
Unidentifiable
7
10.4%
Human labelling
6
9.0%
Table 1: Correctness judges used in surveyed papers.
Figure 2: Posterior F1 (horizontal axis) for a subset of judges for TriviaQA-Llama8b. Rows are sorted by posterior median (circle) and display two high density intervals (50% and 95%; thick/thin line segments).
Figure 3: An example data point of question (Q) and references (R) from TriviaQA and responses by Llama 8B. The greedy response is judged manually and automatically ( e.g. , via exact match, BLEU, BertScore, LLM judges, etc. ); positive judgements are blue/solid, negative are red/dashed. We compute various UQs: some ( e.g. , P(True)) qualify the greedy response specifically (in green), others ( e.g. , entropy) qualify the entire conditional distribution (estimated using samples; in yellow).
Best Quantifier (Rank #1)
Medium Quantifier (Rank #4)
Worst Quantifier (Rank #7)
Model/Dataset pair
Actual
Rank G
Rank M
Rank B
Actual
Rank G
Rank M
Rank B
Actual
Rank G
Rank M
Rank B
TriviaQA/Llama8b
VC
2.3
4.3
4.4
SE
4.2
3.5
3.6
SC
6.8
6.0
5.8
TriviaQA/Llama3b
P(Seq.)
2.2
1.4
2.1
VC
2.5
5.6
5.2
SC
7.0
5.6
6.1
TriviaQA/Qwen7b
SC
1.5
6.0
6.3
N.P(Seq.)
5.0
2.0
2.8
E
6.0
3.3
1.3
TriviaQA/Qwen0.5b
SC
1.0
3.7
5.2
N.P(Seq.)
3.0
4.2
3.5
P(True)
6.0
4.6
4.8
AmbigQA_non/Llama8b
VC
1.4
2.8
4.7
P(True)
3.9
5.8
5.1
E
5.0
4.7
2.2
Table 2: We separate the correctness judges as ‘good’ (G; F1>0.7 ), ‘medium’ (M; 0.3<F1<0.7 ) or ‘bad’ (B; F1<0.3 ) in terms of their F1 score, and identify the best UQ (ranked 1st), medium-quality UQ (ranked 4th) and worst UQ (ranked 7th), according to their oracle AUROC values. We show how the oracle rankings of UQs are affected on average.
Best Quantifier (Rank #1)
Medium Quantifier (Rank #4)
Worst Quantifier (Rank #7)
Model/Dataset pair
Actual
AUROC
Diff. G
Diff. M
Diff. B
Actual
AUROC
Diff. G
Diff. M
Diff. B
Actual
AUROC
Diff. G
Diff. M
Diff. B
TriviaQA/Llama8b
VC
0.92
-0.11
-0.22
-0.19
SE
0.76
-0.03
-0.03
0.02
SC
0.69
-0.07
-0.05
-0.01
TriviaQA/Llama3b
P(Seq.)
0.88
-0.09
-0.10
0.03
VC
0.79
0.00
-0.12
-0.02
SC
0.67
-0.06
-0.02
0.06
TriviaQA/Qwen7b
SC
0.82
-0.08
-0.18
-0.11
N.P(Seq.)
0.72
-0.08
0.08
0.23
E
0.67
-0.05
0.08
0.29
TriviaQA/Qwen0.5b
SC
0.73
0.02
-0.14
0.03
N.P(Seq.)
0.56
0.03
0.05
0.32
P(True)
0.49
0.02
0.07
0.31
AmbigQA_non/Llama8b
VC
0.78
-0.02
-0.07
-0.31
P(True)
0.68
-0.01
-0.10
-0.14
E
0.58
0.02
0.06
0.25
Table 3: We separate the correctness judges as ‘good’ (G; F1>0.7 ), ‘medium’ (M; 0.3<F1<0.7 ) or ‘bad’ (B; F1<0.3 ) in terms of their F1 score and identify the best UQ (ranked 1st), medium-quality UQ (ranked 4th) and worst UQ (ranked 7th), according to their oracle AUROC values. We show how the oracle AUROC values of UQs are affected on average.
Figure 4: F1 scores ( ↑ ) for judges against the RBO score ( ↑ ) of human- vs judge-powered UQ rankings. Even within a narrow range of F1 values ( i.e. , similar-quality judges), there’s considerable variability in their UQ evaluation performance.
Figure 5: Error trade-offs in SP (with threshold 0.5 ) for various UQs under best automated judges (TriviaQA-Llama8B). For a detailed guide on how to interpret this figure, see Section 5.2 .
Figure 6: Error impacts on AUROC for various quantifiers under best automated judges, for TriviaQA-Llama8B. For a detailed guide on how to interpret this figure, see Section 5.2 .
Appendix figures & tables24 assets
Supplementary material from the paper’s appendix.
Appendix
Paper
Judge Type
Judge
Evaluation method
Khanmohammadi et al. (2026)
Exact/fuzzy matching, LLM judge
String match, LLM (GPT-5-nano)
AUROC, AUCPR, ECE, Brier
Testoni and Calixto (2026)
Exact/fuzzy matching, LLM judge, Human validation
String match, LLM (LLaMA-3.1-8B-Instruct), Human validation
AUROC, ECE, Brier
Kunitomo-Jacquin et al. (2026)
LLM judge
LLM (gpt4o-mini)
AUROC, AUARC
Appendix
Table 4: An overview of all tables
UQ Eval. metric
# Papers
Proportion
AUROC
40
59.7%
ECE
26
38.8%
AUCPR
12
17.9%
AURAC
11
16.4%
Other
11
16.4%
Brier
10
14.9%
Appendix
Table 5: UQ evaluation metrics used in retrieved papers.
Figure 7: Top row shows AUROC values, middle row shows differences between oracle and system AUROC values and bottom row rankings for various uncertainty quantifiers (ordered from left to right in a descending order of performance) and correctness functions (ordered from top to bottom in an increasing order of quality).
Figure 8: Top row shows AUROC values, middle row shows differences between oracle and system AUROC values and bottom row rankings for various uncertainty quantifiers (ordered from left to right in a descending order of performance) and correctness functions (ordered from top to bottom in an increasing order of quality).
Figure 9: Top row shows AUROC values, middle row shows differences between oracle and system AUROC values and bottom row rankings for various uncertainty quantifiers (ordered from left to right in a descending order of performance) and correctness functions (ordered from top to bottom in an increasing order of quality).
Figure 10: Top row shows AUROC values, middle row shows differences between oracle and system AUROC values and bottom row rankings for various uncertainty quantifiers (ordered from left to right in a descending order of performance) and correctness functions (ordered from top to bottom in an increasing order of quality).
Figure 11: Top row shows AUROC values, middle row shows differences between oracle and system AUROC values and bottom row rankings for various uncertainty quantifiers (ordered from left to right in a descending order of performance) and correctness functions (ordered from top to bottom in an increasing order of quality).
Figure 12: Top row shows AUROC values, middle row shows differences between oracle and system AUROC values and bottom row rankings for various uncertainty quantifiers (ordered from left to right in a descending order of performance) and correctness functions (ordered from top to bottom in an increasing order of quality).
Figure 13: Left column shows results given manual labels obtained from new annotator, and right column shows results from original annotator.
Figure 14: Left column shows results given manual labels obtained from new annotator, and right column shows results from original annotator.
Figure 15: Top two plots correspond to Llama 3B, TriviaQA. Bottom two plots correspond to Llama8b, TriviaQA.
Figure 16: Top two plots correspond to Qwen 0.5B, TriviaQA. Bottom two plots correspond to Qwen 7B, TriviaQA.
Figure 17: Top two plots correspond to Llama 3B, AmbigQA (unambiguous). Bottom two plots correspond to Llama 8B, AmbigQA (unambiguous).
Figure 18: Top two plots correspond to Qwen 0.5B, AmbigQA (unambiguous). Bottom two plots correspond to Qwen 7B, AmbigQA (unambiguous).
Figure 19: Top two plots correspond to Llama 3B, AmbigQA (ambiguous). Bottom two plots correspond to Llama 8B, AmbigQA (ambiguous).
Figure 20: Top two plots correspond to Qwen 0.5B, AmbigQA (ambiguous). Bottom two plots correspond to Qwen 7B, AmbigQA (ambiguous).
Figure 21: Top row shows risk at fixed coverage values, middle row shows differences and bottom row rankings for various uncertainty quantifiers.
Figure 22: Top row shows risk at fixed coverage values, middle row shows differences and bottom row rankings for various uncertainty quantifiers.
Figure 23: Top row shows Coverage at fixed risk values, middle row shows differences and bottom row rankings for various uncertainty quantifiers.
Figure 24: Top row shows Coverage at fixed risk values, middle row shows differences and bottom row rankings for various uncertainty quantifiers.
Figure 25: F1 distributions for TriviaQA dataset and Llama 8B (top left corner), Llama 3B (top right corner), Qwen 7B (bottom left corner) and Qwen 0.5B (bottom right corner).
Figure 26: F1 distributions for AmbigQA (non-ambiguous) dataset and Llama 8B (top left corner), Llama 3B (top right corner), Qwen 7B (bottom left corner) and Qwen 0.5B (bottom right corner).
Figure 27: F1 distributions for AmbigQA (ambiguous) dataset and Llama 8B (top left corner), Llama 3B (top right corner), Qwen 7B (bottom left corner) and Qwen 0.5B (bottom right corner).
Figure 28: F1 scores of automated judges against RBO scores for all dataset-generator pairs.
Uncertainty Quantification (UQ) is widely regarded as the primary safeguard for deploying Large Language Models (LLMs) in high-stakes domains. However, we argue that the field suffers from a category error: mainstream UQ methods for LLMs are just unsupervised clustering algorithms. We demonstrate that most current approaches inherently quantify the internal consistency of the model's generations rather than their external correctness. Consequently, current methods are fundamentally blind to factual reality and fail to detect ``confident hallucinations,'' where models exhibit high confidence in stable but incorrect answers. Therefore, the current UQ methods may create a deceptive sense of safety when deploying the models with uncertainty. In detail, we identify three critical pathologies resulting from this dependence on internal state: a hyperparameter sensitivity crisis that renders deployment unsafe, an internal evaluation cycle that conflates stability with truth, and a fundamental lack of ground truth that forces reliance on unstable proxy metrics to evaluate uncertainty. To resolve this impasse, we advocate for a paradigm shift to UQ and outline a roadmap for the research community to adopt better evaluation metrics and settings, implement mechanism changes for native uncertainty, and anchor verification in objective truth, ensuring that model confidence serves as a reliable proxy for reality.
Tiejin Chen, Longchao Da, Xiaoou Liu +1
School of Computing and Augmented Intelligence, Arizona State University
LLM-as-a-Judge evaluation has become a standard tool for assessing base model performance. However, characterizing performance via the naive estimator, i.e., raw judge outputs, is systematically biased. Recent work has proposed estimators to correct this bias, but their reliability depends critically on judge quality and, for model comparisons, on calibration stability. Sharing calibration across compared models is practically attractive but can introduce severe bias, including cases where the comparison estimate points in the wrong direction with high apparent confidence. We study these failure modes through analytical results, simulations over judge quality (J) and cross-model calibration instability (ΔJ), and a real-data MMLU-Pro case study with sign reversal. We propose J and ΔJ as diagnostics for when corrected estimates, especially shared-calibration comparisons, are likely unreliable, and provide reporting guidance for LaaJ evaluation.
The task of Error Prediction, namely predicting whether a model output is correct, is commonly tackled with Uncertainty Quantification (UQ). However, while uncertainty metrics capture when models lack knowledge or capacity to make a prediction, they also reflect aleatoric uncertainty, which is inherent in the model input and context. This paper presents a method for improving error prediction for Large Language Models (LLMs), by disentangling input ambiguity from UQ signal. We conduct experiments on the task of Question Answering (QA) with six UQ metrics and show that UQ metrics are more predictive of errors on unambiguous instances than on questions with multiple plausible answers. We use Gated Experts and Selective Prediction to incorporate gold and predicted ambiguity labels into the error prediction pipeline. We find that ambiguity information improves error prediction scores across model families, training and evaluation paradigms, datasets (including allegedly unambiguous ones), and sources of aleatoric uncertainty, yielding improvements of over 10 points of PRR for individual UQ metrics on standard datasets.
Ieva Raminta Staliūnaitė, James Bishop, Andreas Vlachos
University of Cambridge · The Alan Turing Institute