Human-model alignment is critical for trustworthy AI-assisted decision-making systems. Yet, most work evaluates model predictions against single ground-truth labels, overlooking that humans themselves often disagree on labels, a signal of genuine ambiguity. We investigate whether models struggle on the same instances that humans find difficult. We measure this on two vision datasets (FER+ and CIFAR-10H) where multiple human annotations per image capture human disagreement patterns. We evaluate eight pretrained models across three architectures (ResNet, EfficientNet, MobileNetV3) in two parts: first, whether model uncertainty (softmax confidence, entropy) correlates with human disagreement, and second, whether predictive multiplicity measures (inter-model disagreement, Jensen-Shannon divergence) do. We find that it does not: alignment is weak in both dimensions. At the discrete label level, 50.4% of CIFAR-10H images and 33.5% of FER+ images receive multiple valid classifications from humans, while the models converge on only one. These instances represent a critical failure case where humans perceive ambiguity and would request expert review, yet models decide confidently. At the continuous score level, single-model uncertainty correlates weakly with human disagreement (ρ=0.24−−0.55), and predictive multiplicity provides only modest improvement. Widely-used uncertainty quantification methods do not reliably identify instances humans find ambiguous. Model uncertainty should not be treated as a trustworthy signal by default for decision-making in high-stakes scenarios.
Figures & tables
Fig. 1 : Representative FER+ instances for the four human–model disagreement profiles. Each row shows the test image with the distribution of its 10 human votes (red) and of the 8 model predictions (grey) over the emotion classes: Neu(tral), Hap(piness), Sur(prise), Sad(ness), Ang(er), Dis(gust), Fea(r), Con(tempt).
Term
Definition
Human disagreement
Uncertainty arising from disagreement among human annotators when assigning labels to the same instance, as defined in [ 10 , 3 ] (there termed human uncertainty ). It captures the inherent ambiguity or subjectivity of the ground-truth label.
Model uncertainty
Uncertainty in a single model’s prediction, typically quantified using predictive entropy or another uncertainty score derived from its predicted class probabilities [ 12 , 13 ] . It captures how confident a model is about its prediction.
Predictive multiplicity
The disagreement in model predictions from a set of nearly equivalent accuracy models, or the Rashomon set of models [ 11 ] . The extent to which equivalent models reach conflicting predictions for the same instance.
Disagreement profile
The joint certainty status of an instance: whether the annotators and the models are each internally unanimous or divided, yielding the four profiles MC, MoD, HoD, and MD (Sec. III-F).
Alignment
The degree to which model uncertainty or disagreement corresponds to human disagreement. How well model uncertainty or multiplicity reflects the disagreement observed among human annotators.
TABLE I : Definitions of key terms used in this work.
ResNet
EfficientNet
MobileNetV3
18
34
50
B0
B1
B2
Small
Large
CIFAR-10H
94.34
96.31
96.05
96.13
94.96
96.20
92.95
95.10
FER+
77.38
77.58
78.44
78.28
80.32
77.41
77.86
78.72
TABLE II : Test accuracy (%) of the 8 fine-tuned models (three families) on CIFAR-10H and FER+ .
Models Agree
Models Disagree
Total
Humans Agree
MC: 815 ( 22.8 %)
MoD: 186 ( 5.2 %)
1001 ( 28.0 %)
Humans Disagree
HoD: 1195 ( 33.5 %)
MD: 1376 ( 38.5 %)
2571 ( 72.0 %)
Total
2010 ( 56.3 %)
1562 ( 43.7 %)
3572
TABLE III : Human-model agreement on the FER+ test set ( n=3572 ). Cells show counts with percentages of the test set in parentheses.
Models Agree
Models Disagree
Total
Humans Agree
MC: 3310 ( 33.1 %)
MoD: 255 ( 2.6 %)
3565 (35.7%)
Humans Disagree
HoD: 5044 ( 50.4 %)
MD: 1391 ( 13.9 %)
6435 (64.4%)
Total
8354 ( 83.5 %)
1646 ( 16.5 %)
10,000
TABLE IV : Human–model agreement on the CIFAR-10H test set ( n=10,000 ). Cells show counts with percentages of the test set in parentheses.
Fig. 4 : Per-instance human disagreement ( Hhumannorm ) vs. model disagreement for each dataset and model-side measure; colour gives image counts on a log scale. Even in the best-aligned setting ( FER+ , average predictive entropy ( Hˉ ), ρ=0.572 ), instances are widely dispersed off the diagonal. On CIFAR-10H the weak alignment holds for the continuous measure as well ( ρ=0.363 ). For inter-model vote entropy ( Hvotenorm , ρ=0.315 ), the dense band at zero, which is the 83.5% of images on which all eight models agree, spans the full range of human disagreement.
ResNet
EfficientNet
MobileNetV3
18
34
50
B0
B1
B2
Small
Large
CIFAR-10H
SC
-0.276
-0.237
-0.243
-0.340
-0.344
-0.317
-0.344
-0.350
PE
0.275
0.235
0.241
0.344
0.346
0.327
0.346
0.347
FER+
SC
-0.470
-0.450
-0.420
-0.438
-0.550
-0.413
-0.488
-0.451
PE
0.474
0.450
0.421
0.442
0.554
0.416
0.491
0.453
TABLE V : Spearman correlations ( ρ ) between single-model uncertainty measures and normalised human vote entropy on FER+ and CIFAR-10H . All correlations are statistically significant ( p<0.05 ). SC: Softmax Confidence; PE: Predictive Entropy.
Cˉ
Hˉ
Hvotenorm
DJSD
Dwithin
Dbetween
R
CIFAR-10H
−0.360
0.363
0.315
0.184
0.201
0.174
−0.184
FER+
−0.561
0.572
0.421
0.507
0.523
0.471
−0.314
TABLE VI : Spearman correlations ( ρ ) between the aggregate model-side measures (Secs. III-D - III-F ) and human disagreement Hhumannorm . All correlations are statistically significant ( p<0.05 ).
Heriot-Watt University, Edinburgh, Scotland · aequa-tech, Torino, Italy · Laboratory for the Modeling of Biological and Socio-technical Systems, Northeastern University, Boston, MA, USA +2