The wide adoption of LLMs across broad NLG applications heightens the importance of providing users with the means to avert errors and hallucinations. Uncertainty quantification is poised to fill that gap; with low uncertainty (high confidence), as a proxy for correctness, allowing users to be selective (e.g., reject low-confidence, likely incorrect responses). Correlation between confidence and correctness then serves as a useful criterion for evaluation of uncertainty quantifiers (UQs). But in NLG, where diverse responses can be adequate to a prompt, obtaining reliable correctness judgements is not simple, especially without human intervention. Errors in automated judgement are hardly avoidable and known to diminish the reliability of evaluation protocols (Santilli et al., 2025; Ielanskyi et al., 2025). In a meta-analysis of published work, we show that automated judgement is the present norm. Besides, automated judgements are rarely validated against human ones, and the validation of the UQ evaluation they automate is even rarer. With experiments in question answering, using 4 LLMs, human and automated judgements and 7 popular UQs, we find that i) a judge's performance can only coarsely predict the observed impact of its errors on the reliability of UQ evaluation, and that ii) judgement errors tend to misrepresent informative UQs most. We link these observations to patterns of correlation between confidence and categories of judgement error.
Figures & tables
Figure 1: An informative UQ (top-left) assigns higher confidence to responses that humans deem correct (C) than to those that humans deem wrong (W), allowing a decision maker to be selective ( e.g. , by rejecting lower-confidence responses wrt a threshold (dashed line)). If assigned confidences do not neatly separate Cs from Ws (top-right), the UQ leads to selection errors (shown in red). The degree of separation gives us a criterion to choose amongst UQs. But when judgement is automated (bottom), as routinely done in UQ evaluation, judgement errors (C → W or W → C; shown as diamonds) distort our view of the relative merits of the UQs. This affects informative UQs most: on the left, we go from 0/6 to 3/6 selection errors (2 rejected Cs and 1 accepted W), while on the right we go from 3/6 to 4/6 errors.
Correctness judge
# Papers
Proportion
LLM judge
26
38.8%
Overlap based
19
28.3%
Exact/fuzzy match
14
20.9%
Semantic similarity
11
16.4%
Unidentifiable
7
10.4%
Human labelling
6
9.0%
Table 1: Correctness judges used in surveyed papers.
Figure 2: Posterior F1 (horizontal axis) for a subset of judges for TriviaQA-Llama8b. Rows are sorted by posterior median (circle) and display two high density intervals (50% and 95%; thick/thin line segments).
Figure 3: An example data point of question (Q) and references (R) from TriviaQA and responses by Llama 8B. The greedy response is judged manually and automatically ( e.g. , via exact match, BLEU, BertScore, LLM judges, etc. ); positive judgements are blue/solid, negative are red/dashed. We compute various UQs: some ( e.g. , P(True)) qualify the greedy response specifically (in green), others ( e.g. , entropy) qualify the entire conditional distribution (estimated using samples; in yellow).
Best Quantifier (Rank #1)
Medium Quantifier (Rank #4)
Worst Quantifier (Rank #7)
Model/Dataset pair
Actual
Rank G
Rank M
Rank B
Actual
Rank G
Rank M
Rank B
Actual
Rank G
Rank M
Rank B
TriviaQA/Llama8b
VC
2.3
4.3
4.4
SE
4.2
3.5
3.6
SC
6.8
6.0
5.8
TriviaQA/Llama3b
P(Seq.)
2.2
1.4
2.1
VC
2.5
5.6
5.2
SC
7.0
5.6
6.1
TriviaQA/Qwen7b
SC
1.5
6.0
6.3
N.P(Seq.)
5.0
2.0
2.8
E
6.0
3.3
1.3
TriviaQA/Qwen0.5b
SC
1.0
3.7
5.2
N.P(Seq.)
3.0
4.2
3.5
P(True)
6.0
4.6
4.8
AmbigQA_non/Llama8b
VC
1.4
2.8
4.7
P(True)
3.9
5.8
5.1
E
5.0
4.7
2.2
Table 2: We separate the correctness judges as ‘good’ (G; F1>0.7 ), ‘medium’ (M; 0.3<F1<0.7 ) or ‘bad’ (B; F1<0.3 ) in terms of their F1 score, and identify the best UQ (ranked 1st), medium-quality UQ (ranked 4th) and worst UQ (ranked 7th), according to their oracle AUROC values. We show how the oracle rankings of UQs are affected on average.
Best Quantifier (Rank #1)
Medium Quantifier (Rank #4)
Worst Quantifier (Rank #7)
Model/Dataset pair
Actual
AUROC
Diff. G
Diff. M
Diff. B
Actual
AUROC
Diff. G
Diff. M
Diff. B
Actual
AUROC
Diff. G
Diff. M
Diff. B
TriviaQA/Llama8b
VC
0.92
-0.11
-0.22
-0.19
SE
0.76
-0.03
-0.03
0.02
SC
0.69
-0.07
-0.05
-0.01
TriviaQA/Llama3b
P(Seq.)
0.88
-0.09
-0.10
0.03
VC
0.79
0.00
-0.12
-0.02
SC
0.67
-0.06
-0.02
0.06
TriviaQA/Qwen7b
SC
0.82
-0.08
-0.18
-0.11
N.P(Seq.)
0.72
-0.08
0.08
0.23
E
0.67
-0.05
0.08
0.29
TriviaQA/Qwen0.5b
SC
0.73
0.02
-0.14
0.03
N.P(Seq.)
0.56
0.03
0.05
0.32
P(True)
0.49
0.02
0.07
0.31
AmbigQA_non/Llama8b
VC
0.78
-0.02
-0.07
-0.31
P(True)
0.68
-0.01
-0.10
-0.14
E
0.58
0.02
0.06
0.25
Table 3: We separate the correctness judges as ‘good’ (G; F1>0.7 ), ‘medium’ (M; 0.3<F1<0.7 ) or ‘bad’ (B; F1<0.3 ) in terms of their F1 score and identify the best UQ (ranked 1st), medium-quality UQ (ranked 4th) and worst UQ (ranked 7th), according to their oracle AUROC values. We show how the oracle AUROC values of UQs are affected on average.
Figure 4: F1 scores ( ↑ ) for judges against the RBO score ( ↑ ) of human- vs judge-powered UQ rankings. Even within a narrow range of F1 values ( i.e. , similar-quality judges), there’s considerable variability in their UQ evaluation performance.
Figure 5: Error trade-offs in SP (with threshold 0.5 ) for various UQs under best automated judges (TriviaQA-Llama8B). For a detailed guide on how to interpret this figure, see Section 5.2 .
Figure 6: Error impacts on AUROC for various quantifiers under best automated judges, for TriviaQA-Llama8B. For a detailed guide on how to interpret this figure, see Section 5.2 .
Appendix figures & tables24 assets
Supplementary material from the paper’s appendix.
Appendix
Paper
Judge Type
Judge
Evaluation method
Khanmohammadi et al. (2026)
Exact/fuzzy matching, LLM judge
String match, LLM (GPT-5-nano)
AUROC, AUCPR, ECE, Brier
Testoni and Calixto (2026)
Exact/fuzzy matching, LLM judge, Human validation
String match, LLM (LLaMA-3.1-8B-Instruct), Human validation
AUROC, ECE, Brier
Kunitomo-Jacquin et al. (2026)
LLM judge
LLM (gpt4o-mini)
AUROC, AUARC
Appendix
Table 4: An overview of all tables
UQ Eval. metric
# Papers
Proportion
AUROC
40
59.7%
ECE
26
38.8%
AUCPR
12
17.9%
AURAC
11
16.4%
Other
11
16.4%
Brier
10
14.9%
Appendix
Table 5: UQ evaluation metrics used in retrieved papers.
Figure 7: Top row shows AUROC values, middle row shows differences between oracle and system AUROC values and bottom row rankings for various uncertainty quantifiers (ordered from left to right in a descending order of performance) and correctness functions (ordered from top to bottom in an increasing order of quality).
Figure 8: Top row shows AUROC values, middle row shows differences between oracle and system AUROC values and bottom row rankings for various uncertainty quantifiers (ordered from left to right in a descending order of performance) and correctness functions (ordered from top to bottom in an increasing order of quality).
Figure 9: Top row shows AUROC values, middle row shows differences between oracle and system AUROC values and bottom row rankings for various uncertainty quantifiers (ordered from left to right in a descending order of performance) and correctness functions (ordered from top to bottom in an increasing order of quality).
Figure 10: Top row shows AUROC values, middle row shows differences between oracle and system AUROC values and bottom row rankings for various uncertainty quantifiers (ordered from left to right in a descending order of performance) and correctness functions (ordered from top to bottom in an increasing order of quality).
Figure 11: Top row shows AUROC values, middle row shows differences between oracle and system AUROC values and bottom row rankings for various uncertainty quantifiers (ordered from left to right in a descending order of performance) and correctness functions (ordered from top to bottom in an increasing order of quality).
Figure 12: Top row shows AUROC values, middle row shows differences between oracle and system AUROC values and bottom row rankings for various uncertainty quantifiers (ordered from left to right in a descending order of performance) and correctness functions (ordered from top to bottom in an increasing order of quality).
Figure 13: Left column shows results given manual labels obtained from new annotator, and right column shows results from original annotator.
Figure 14: Left column shows results given manual labels obtained from new annotator, and right column shows results from original annotator.
Figure 15: Top two plots correspond to Llama 3B, TriviaQA. Bottom two plots correspond to Llama8b, TriviaQA.
Figure 16: Top two plots correspond to Qwen 0.5B, TriviaQA. Bottom two plots correspond to Qwen 7B, TriviaQA.
Figure 17: Top two plots correspond to Llama 3B, AmbigQA (unambiguous). Bottom two plots correspond to Llama 8B, AmbigQA (unambiguous).
Figure 18: Top two plots correspond to Qwen 0.5B, AmbigQA (unambiguous). Bottom two plots correspond to Qwen 7B, AmbigQA (unambiguous).
Figure 19: Top two plots correspond to Llama 3B, AmbigQA (ambiguous). Bottom two plots correspond to Llama 8B, AmbigQA (ambiguous).
Figure 20: Top two plots correspond to Qwen 0.5B, AmbigQA (ambiguous). Bottom two plots correspond to Qwen 7B, AmbigQA (ambiguous).
Figure 21: Top row shows risk at fixed coverage values, middle row shows differences and bottom row rankings for various uncertainty quantifiers.
Figure 22: Top row shows risk at fixed coverage values, middle row shows differences and bottom row rankings for various uncertainty quantifiers.
Figure 23: Top row shows Coverage at fixed risk values, middle row shows differences and bottom row rankings for various uncertainty quantifiers.
Figure 24: Top row shows Coverage at fixed risk values, middle row shows differences and bottom row rankings for various uncertainty quantifiers.
Figure 25: F1 distributions for TriviaQA dataset and Llama 8B (top left corner), Llama 3B (top right corner), Qwen 7B (bottom left corner) and Qwen 0.5B (bottom right corner).
Figure 26: F1 distributions for AmbigQA (non-ambiguous) dataset and Llama 8B (top left corner), Llama 3B (top right corner), Qwen 7B (bottom left corner) and Qwen 0.5B (bottom right corner).
Figure 27: F1 distributions for AmbigQA (ambiguous) dataset and Llama 8B (top left corner), Llama 3B (top right corner), Qwen 7B (bottom left corner) and Qwen 0.5B (bottom right corner).
Figure 28: F1 scores of automated judges against RBO scores for all dataset-generator pairs.