cs.AISep 28, 2026

Does Model Uncertainty Track Human Ambiguity? Evidence from Multi-Annotator Vision Benchmarks

Authors: Manya Singh, Arjun Pakrashi

Organizations: School of Computer Science, University College Dublin, Ireland

Abstract

Human-model alignment is critical for trustworthy AI-assisted decision-making systems. Yet, most work evaluates model predictions against single ground-truth labels, overlooking that humans themselves often disagree on labels, a signal of genuine ambiguity. We investigate whether models struggle on the same instances that humans find difficult. We measure this on two vision datasets (FER+ and CIFAR-10H) where multiple human annotations per image capture human disagreement patterns. We evaluate eight pretrained models across three architectures (ResNet, EfficientNet, MobileNetV3) in two parts: first, whether model uncertainty (softmax confidence, entropy) correlates with human disagreement, and second, whether predictive multiplicity measures (inter-model disagreement, Jensen-Shannon divergence) do. We find that it does not: alignment is weak in both dimensions. At the discrete label level, 50.4% of CIFAR-10H images and 33.5% of FER+ images receive multiple valid classifications from humans, while the models converge on only one. These instances represent a critical failure case where humans perceive ambiguity and would request expert review, yet models decide confidently. At the continuous score level, single-model uncertainty correlates weakly with human disagreement (ρ=0.24−−0.55ρ= 0.24--0.55), and predictive multiplicity provides only modest improvement. Widely-used uncertainty quantification methods do not reliably identify instances humans find ambiguous. Model uncertainty should not be treated as a trustworthy signal by default for decision-making in high-stakes scenarios.

Figures & tables

Explore similar work

CardsList
  1. From Model Uncertainty to Human Attention: Localization-Aware Visual Cues for Scalable Annotation Review

    May 12, 2026Moussa Kassem Sbeyti, Joshua Holstein, Philipp Spitzer +2Automatic LabellingAnnotations

  2. An Assessment of Human vs. Model Uncertainty in Soft-Label Learning and Calibration

    May 18, 2026Maja Pavlovic, Silviu Paun, Massimo PoesioArtificial Intelligence Alignment

  3. The Ghost Annotator: a Framework to Explore Human Label Variation in Content Moderation through Conformal Prediction

    Jun 1, 2026Mirko Lai, Alessandra Urbinati, Simona Frenda +2Large Language Model AnnotationsHuman Annotators