Semantic uncertainty quantification for large language models rests on a common template: sample several answers, measure how much they agree, and treat disagreement as uncertainty. We first formalize this template as two separate roles: an operator that compares two answers, and an aggregator that combines all pairwise comparisons into a scalar. Existing methods differ almost entirely in how they aggregate, while taking the operator off the shelf, typically an NLI model or a generic sentence encoder. We show that this reliance on off-the-shelf operators is the primary bottleneck of semantic UQ: they do not accurately measure factual equivalence of multiple answers to the same question. We resolve this with a deliberately simple recipe: a single encoder trained contrastively to isolate the targeted fact, utilizing synthetic data generated by an LLM and dataset both disjoint from all evaluation settings. Integrating the resulting operator into existing methods improves performance on 120 of 126 evaluation settings (95%) spanning 18 model dataset combinations across language and vision-language models. The best variant reaches 0.76 mean AUROC against 0.68 for the strongest baseline, while replacing the quadratic cross-encoder comparisons of entailment-based operators with one encoder pass per answer. The uniformity of the improvement supports the view that the operator, not the aggregator, is the limiting factor. The same operator also improves single generation token-level estimators: the norm it assigns to each token measures how much that token bears on the answer, and reweighting token log-likelihoods accordingly sharpens the estimate.
Figures & tables
Figure 1: The Operator Bottleneck in Semantic UQ. Semantic uncertainty methods decompose into a pairwise comparison operator ( s ) and a structural aggregator ( F ). While prior work focuses on designing sophisticated aggregators, they inherit off-the-shelf operators that suffer from misaligned training objectives or severe scaling bottlenecks. We identify the operator as the true bottleneck and introduce a contrastively trained, question-conditioned bi-encoder. Serving as a universal drop-in replacement, our learned operator universally improves the performance of existing aggregators at a fraction of the computational cost of NLI baselines.
Figure 2: Learning Semantic Geometry from Predictive Variability. (A) Mining Predictive Variability (Section 4.2 ): We extract positive and negative semantic triplets by evaluating repeated LLM generations for factual correctness, naturally capturing the model’s native syntax distribution. (B) Contrastive Factual Equivalence (Sections 4.1 and 4.3 ): A shared bi-encoder contextualizes the answers against the query. The contrastive margin separates hard negatives (incorrect answers to the same question) from valid paraphrases. (C) Zero-Shot Inference & UQ (Section 5 ): The frozen encoder generates symmetric O(N) comparison matrices that drop seamlessly into existing aggregators, while its geometric norm enables zero-cost token reweighting.
ADVQA
OKVQA
VizWiz
HotpotQA
TriviaQA
WebQuestions
post-hoc
Method
AUROC
ECE
AUROC
ECE
AUROC
ECE
AUROC
ECE
AUROC
ECE
AUROC
ECE
Δ
compute
CAE
.637
.116
.643
.229
.647
.199
.673
.235
.583
.212
.490
.176
1.11 s
CAEEnc
.663
.081
.733
.194
.674
.159
.722
.215
.840
.135
.716
.055
+.113
0.03 s
SE
.637
.114
.643
.228
.644
.198
.674
.235
.577
.212
.488
.175
1.18 s
SEEnc
.666
.080
.733
.193
.675
.159
.725
.214
.840
.133
.719
.058
+.116
0.03 s
KLE-Heat
.675
.080
.678
.187
.690
.161
.719
.208
.747
.127
.555
.112
1.21 s
Table 1: Main results. AUROC ( ↑ ) and ECE ( ↓ ) for each estimator in its published form ( base ) and with our operator substituted ( shaded ), averaged over the 18 model–dataset combinations and reported by dataset (top) and by model (bottom). Δ is the change in mean AUROC over all columns.
Appendix figures & tables9 assets
Supplementary material from the paper’s appendix.
Appendix
Correctness
Answer 1
Answer 2
Similarity
Both correct
The capital of France is Paris.
Paris is the capital of France.
.975
Both incorrect
The capital of France is Madrid.
Madrid is the capital of France.
.980
Appendix
Table B.1: Similarity assigned by our encoder to pairs of semantically equivalent answers, for correct and incorrect content. Question: “What is the capital of France?”
Backbone
AUROC
sentence-t5-large
.759
all-roberta-large-v1
.759
all-mpnet-base-v2
.756
multi-qa-mpnet-base-dot-v1
.755
bge-base-en-v1.5
.752
gtr-t5-large
.750
Appendix
Table B.2: The framework is insensitive to the backbone. Mean AUROC of COSEnc over the 18 model–dataset combinations, for six sentence encoders used as the starting point of the training procedure of Sec. 4 , with all other settings held fixed.
ADVQA
OKVQA
VizWiz
HotpotQA
TriviaQA
WebQuestions
Mean
hard negatives (ours)
.694
.758
.741
.765
.852
.741
.759
random negatives
.666
.690
.697
.738
.823
.618
.705
untrained encoder
.651
.673
.705
.707
.785
.554
.679
no negatives
.585
.616
.610
.687
.760
.640
.650
idefics2-8B
Qwen2.5-VL-7B
llava-1.5-7B
Qwen2.5-7B
Phi-3.5-mini
Mistral-7B
Mean
hard negatives (ours)
.730
.740
.723
.790
.793
.775
.759
Appendix
Table B.3: Hard negatives matter, not just negatives. AUROC of COSEnc for the untrained encoder and under three training regimes: no negatives (cosine-similarity loss on positive pairs only), random negatives (drawn from a different question), and hard negatives (drawn from the same question, used throughout this paper), averaged over models (top) and over datasets (bottom).
Margin
Learning rate
AUROC
0.10
2×10−5
.765
0.15
2×10−5
.758
0.20
2×10−5
.759
0.30
2×10−5
.757
0.50
2×10−5
.749
0.20
1×10−5
.759
Appendix
Table B.4: Small variations across hyperparameters. AUROC of COSEnc , averaged over the 18 model–dataset combinations, when the margin and learning rate are varied around the configuration used throughout (shaded). All other settings are held fixed.
AUROC
ECE
τ=0.3
.656
.139
τ=0.4
.693
.138
τ=0.5
.724
.140
τ=0.6
.743
.137
τ=0.7
.751
.141
Appendix
Table B.5: Sensitivity to the clustering threshold. AUROC and ECE of CAEEnc over the 18 model–dataset combinations, as the threshold τ applied to SE is varied. The value used throughout the paper is shaded.
Qwen2.5-7B
Phi-3.5-mini
Mistral-7B
Method
AUROC
ECE
AUROC
ECE
AUROC
ECE
HotpotQA
CAE
.683 ±.070
.234 ±.050
.651 ±.021
.259 ±.014
.685 ±.016
.212 ±.012
CAEEnc
.724 ±.067
.216 ±.051
.722 ±.019
.235 ±.014
.719 ±.015
.193 ±.012
SE
.684 ±.067
.234 ±.046
.653 ±.020
.258 ±.014
.686 ±.015
.212 ±.013
SEEnc
.726 ±.064
.215 ±.050
.726 ±.021
.235 ±.014
.722 ±.016
.193 ±.012
Appendix
Table G.1: Per-cell results on the textual benchmarks. AUROC ( ↑ ) and ECE ( ↓ ) with bootstrapped 95% confidence intervals, for each estimator in its published form ( base ) and with our operator substituted ( shaded ), reported for each of the three text-only models.
ADVQA
OKVQA
VizWiz
HotpotQA
TriviaQA
WebQuestions
Mean
with question
.694
.758
.741
.765
.852
.741
.759
answer only
.668
.735
.730
.759
.842
.719
.742
idefics2-8B
Qwen2.5-VL-7B
llava-1.5-7B
Qwen2.5-7B
Phi-3.5-mini
Mistral-7B
Mean
with question
.730
.740
.723
.790
.793
.775
.759
answer only
.711
.725
.697
.776
.782
.761
.742
Appendix
Table G.2: Conditioning on the question helps in every setting. AUROC of COSEnc with and without the question, by dataset (top) and by model (bottom). In the answer-only variant, the question is removed from both training and evaluation.
Greedy
Majority vote
Method
AUROC
Δ
AUROC
Δ
Order preserved
CAE
.612
.629
CAEEnc
.724
+.112
.759
+.130
✓
SE
.610
.626
SEEnc
.726
+.116
.761
+.135
✓
KLE-Heat
.677
.706
Appendix
Table G.3: AUROC under greedy and majority-vote correctness labels. Scores are averaged over the 18 model–dataset combinations. Δ is the AUROC gain of our variant over its base counterpart (over COS-ext for the COS block). The last column indicates whether the ranking within the block is the same under both labelling strategies. Best value per block in bold.
idefics2-8B
Qwen2.5-VL-7B
llava-1.5-7B
Method
AUROC
ECE
AUROC
ECE
AUROC
ECE
ADVQA
CAE
.644 ±.021
.115 ±.018
.631 ±.024
.127 ±.020
.635 ±.022
.107 ±.019
CAEEnc
.649 ±.022
.085 ±.019
.698 ±.021
.059 ±.019
.642 ±.022
.099 ±.020
SE
.646 ±.022
.111 ±.019
.629 ±.025
.129 ±.020
.636 ±.022
.103 ±.018
SEEnc
.652 ±.023
.083 ±.020
.702 ±.022
.060 ±.019
.645 ±.023
.098 ±.020
Appendix
Table G.4: Per-cell results on the visual benchmarks. AUROC ( ↑ ) and ECE ( ↓ ) with bootstrapped 95% confidence intervals, for each estimator in its published form ( base ) and with our operator substituted ( shaded ), reported for each of the three vision-language models.
Institute of Intelligent Software, Guangzhou, China · Institute of Software, Chinese Academy of Sciences, Beijing, China · University of Liverpool, Liverpool, UK