Human-model alignment is critical for trustworthy AI-assisted decision-making systems. Yet, most work evaluates model predictions against single ground-truth labels, overlooking that humans themselves often disagree on labels, a signal of genuine ambiguity. We investigate whether models struggle on the same instances that humans find difficult. We measure this on two vision datasets (FER+ and CIFAR-10H) where multiple human annotations per image capture human disagreement patterns. We evaluate eight pretrained models across three architectures (ResNet, EfficientNet, MobileNetV3) in two parts: first, whether model uncertainty (softmax confidence, entropy) correlates with human disagreement, and second, whether predictive multiplicity measures (inter-model disagreement, Jensen-Shannon divergence) do. We find that it does not: alignment is weak in both dimensions. At the discrete label level, 50.4% of CIFAR-10H images and 33.5% of FER+ images receive multiple valid classifications from humans, while the models converge on only one. These instances represent a critical failure case where humans perceive ambiguity and would request expert review, yet models decide confidently. At the continuous score level, single-model uncertainty correlates weakly with human disagreement (ρ=0.24−−0.55), and predictive multiplicity provides only modest improvement. Widely-used uncertainty quantification methods do not reliably identify instances humans find ambiguous. Model uncertainty should not be treated as a trustworthy signal by default for decision-making in high-stakes scenarios.
Figures & tables
Fig. 1 : Representative FER+ instances for the four human–model disagreement profiles. Each row shows the test image with the distribution of its 10 human votes (red) and of the 8 model predictions (grey) over the emotion classes: Neu(tral), Hap(piness), Sur(prise), Sad(ness), Ang(er), Dis(gust), Fea(r), Con(tempt).
Term
Definition
Human disagreement
Uncertainty arising from disagreement among human annotators when assigning labels to the same instance, as defined in [ 10 , 3 ] (there termed human uncertainty ). It captures the inherent ambiguity or subjectivity of the ground-truth label.
Model uncertainty
Uncertainty in a single model’s prediction, typically quantified using predictive entropy or another uncertainty score derived from its predicted class probabilities [ 12 , 13 ] . It captures how confident a model is about its prediction.
Predictive multiplicity
The disagreement in model predictions from a set of nearly equivalent accuracy models, or the Rashomon set of models [ 11 ] . The extent to which equivalent models reach conflicting predictions for the same instance.
Disagreement profile
The joint certainty status of an instance: whether the annotators and the models are each internally unanimous or divided, yielding the four profiles MC, MoD, HoD, and MD (Sec. III-F).
Alignment
The degree to which model uncertainty or disagreement corresponds to human disagreement. How well model uncertainty or multiplicity reflects the disagreement observed among human annotators.
TABLE I : Definitions of key terms used in this work.
ResNet
EfficientNet
MobileNetV3
18
34
50
B0
B1
B2
Small
Large
CIFAR-10H
94.34
96.31
96.05
96.13
94.96
96.20
92.95
95.10
FER+
77.38
77.58
78.44
78.28
80.32
77.41
77.86
78.72
TABLE II : Test accuracy (%) of the 8 fine-tuned models (three families) on CIFAR-10H and FER+ .
Models Agree
Models Disagree
Total
Humans Agree
MC: 815 ( 22.8 %)
MoD: 186 ( 5.2 %)
1001 ( 28.0 %)
Humans Disagree
HoD: 1195 ( 33.5 %)
MD: 1376 ( 38.5 %)
2571 ( 72.0 %)
Total
2010 ( 56.3 %)
1562 ( 43.7 %)
3572
TABLE III : Human-model agreement on the FER+ test set ( n=3572 ). Cells show counts with percentages of the test set in parentheses.
Models Agree
Models Disagree
Total
Humans Agree
MC: 3310 ( 33.1 %)
MoD: 255 ( 2.6 %)
3565 (35.7%)
Humans Disagree
HoD: 5044 ( 50.4 %)
MD: 1391 ( 13.9 %)
6435 (64.4%)
Total
8354 ( 83.5 %)
1646 ( 16.5 %)
10,000
TABLE IV : Human–model agreement on the CIFAR-10H test set ( n=10,000 ). Cells show counts with percentages of the test set in parentheses.
Fig. 4 : Per-instance human disagreement ( Hhumannorm ) vs. model disagreement for each dataset and model-side measure; colour gives image counts on a log scale. Even in the best-aligned setting ( FER+ , average predictive entropy ( Hˉ ), ρ=0.572 ), instances are widely dispersed off the diagonal. On CIFAR-10H the weak alignment holds for the continuous measure as well ( ρ=0.363 ). For inter-model vote entropy ( Hvotenorm , ρ=0.315 ), the dense band at zero, which is the 83.5% of images on which all eight models agree, spans the full range of human disagreement.
ResNet
EfficientNet
MobileNetV3
18
34
50
B0
B1
B2
Small
Large
CIFAR-10H
SC
-0.276
-0.237
-0.243
-0.340
-0.344
-0.317
-0.344
-0.350
PE
0.275
0.235
0.241
0.344
0.346
0.327
0.346
0.347
FER+
SC
-0.470
-0.450
-0.420
-0.438
-0.550
-0.413
-0.488
-0.451
PE
0.474
0.450
0.421
0.442
0.554
0.416
0.491
0.453
TABLE V : Spearman correlations ( ρ ) between single-model uncertainty measures and normalised human vote entropy on FER+ and CIFAR-10H . All correlations are statistically significant ( p<0.05 ). SC: Softmax Confidence; PE: Predictive Entropy.
Cˉ
Hˉ
Hvotenorm
DJSD
Dwithin
Dbetween
R
CIFAR-10H
−0.360
0.363
0.315
0.184
0.201
0.174
−0.184
FER+
−0.561
0.572
0.421
0.507
0.523
0.471
−0.314
TABLE VI : Spearman correlations ( ρ ) between the aggregate model-side measures (Secs. III-D - III-F ) and human disagreement Hhumannorm . All correlations are statistically significant ( p<0.05 ).
High-quality labeled data is essential for training robust machine learning models, yet obtaining annotations at scale remains expensive. AI-assisted annotation has therefore become standard in large-scale labeling workflows. However, in tasks where model predictions carry two independent components, a class label and spatial boundaries, a model may classify an object with high confidence while mislocalizing it. Existing AI-assisted workflows offer annotators no signal about where spatial errors are most likely. Without such guidance, humans may systematically underinspect subtly misplaced boxes. We address this by studying the effect of visualizing spatial uncertainty via a purpose-built interface. In a controlled study with 120 participants, those receiving uncertainty cues achieve higher label quality while being faster overall. A box-level analysis confirms that the cues redirect annotator effort toward high-uncertainty predictions and away from well-localized boxes. These findings establish localization uncertainty as a lever to improve human-in-the-loop annotation. Code is available at https://mos-ks.github.io/MUHA/.
Moussa Kassem Sbeyti, Joshua Holstein, Philipp Spitzer +2
1Scientific Computing Center, Karlsruhe Institute of Technology. · Institute for Information Systems, Karlsruhe Institute of Technology.
Central to human-aligned AI is understanding the benefits of human-elicited labels over synthetic alternatives. While human soft-labels improve calibration by capturing uncertainty, prior studies conflate these benefits with the implicit correction of mislabeled data (mode shifts), obscuring true effects of soft-labels. We present a controlled audit of soft-label learning across MNIST and a synthetic variant, re-annotating subsets to extract human uncertainty. By decoupling soft-label supervision from underlying label mode shifts, we show that while human soft-labels do provide accuracy gains, their larger value lies in acting as a regularizer that improves model calibration on difficult samples and promotes stable convergence across training runs. Dataset cartography reveals models trained on human soft-labels mirror human uncertainty, whereas those trained on synthetic labels fail to align with humans. Broadly, this work provides a diagnostic testbed for human-AI uncertainty alignment.
Maja Pavlovic, Silviu Paun, Massimo Poesio
Queen Mary University London · Amazon · Queen Mary University London - University of Utrecht
Current research primarily focuses on model performance, while comparatively less attention has been devoted to uncertainty estimation, particularly in settings where LLMs are increasingly used to generate annotated data. We introduce a framework combining conformal prediction with Collaborative Filtering-style annotators' representation to model LLM behavior in relation to human annotators and to analyze patterns of agreement and disagreement. Using Non-Conformity Scores, we introduce the Ghost Prediction metric and the Ghost Annotator representation to quantify cases in which model predictions diverge from all available human annotations. We compute cosine similarity measures to explore differences in model behavior across sociodemographic axes. We evaluated four LLMs of different size and families across four content moderation datasets. Our finding shows that while we find that all models uncertainty increases with annotator disagreement, larger models tend to be more confident in the classification of texts that are not aligned with any human annotation. Finally, the Ghost Annotator framework reveals a consistent and robust pattern of demographic misalignment, suggesting a structural bias likely rooted in pretraining corpora.
Mirko Lai, Alessandra Urbinati, Simona Frenda +2
Heriot-Watt University, Edinburgh, Scotland · aequa-tech, Torino, Italy · Laboratory for the Modeling of Biological and Socio-technical Systems, Northeastern University, Boston, MA, USA +2