Foundation models are increasingly adopted across a wide range of applications, often serving as core blocks within AI systems. Yet different foundation models may encode the same input from multiple different views, leading to substantial representation disagreement, which we term Rashomon Representation. Such disagreement often signals inputs that a given model encodes in a way inconsistent with other models, offering a valuable yet underexplored signal for input reliability estimation. While prior work has largely focused on measuring disagreement across multiple models with a representation set, we instead focus on predicting disagreement from a single representation. We hypothesize that this disagreement follows some consistent, input-dependent patterns rather than occurring at random. To test this, we quantify disagreement by comparing each sample's nearest neighbors across different models' representation spaces, then train a lightweight predictor that estimates disagreement from a single model's representation. At inference time, given a new input, the predictor uses that input's representation to tell whether it aligns with or diverges from those of other models. Extensive experiments across diverse foundation models and datasets show that representational disagreement is indeed input-dependent, predictable, and generalizable, enabling efficient reliability estimation of foundation models.
Figures & tables
Figure 1 : Illustration of Rashomon Representation : different foundation models can produce distinct representations of the same input, reflecting complementary views of its underlying information.
Figure 2 : Training framework for predicting inter-model representational disagreement. We measure neighborhood agreement between an anchor model and comparison models over a shared reference set to obtain the supervision signal y(x) . A lightweight predictor gϕ learns to approximate y(x) using only the anchor-model representation by minimizing L(ϕ) . All foundation models remain frozen, and only the predictor parameters are optimized.
Figure 3 : Measured and predicted NAk across three encoders. Hexagons indicate sample density, while markers and error bars show binned means and 90% intervals. The strong agreement between measured and predicted values demonstrates that cross-model representational disagreement can be reliably inferred from single-model representations.
Comparison models
Anchor
Pearson r↑
Swin-B [-1pt] SigLIP 2 [-1pt] ConvNeXt-B
CLIP
0.734
DINOv2
0.751
ViT-21K
0.701
CLIP [-1pt] DINOv2 [-1pt] ViT-21K
Swin-B
0.770
SigLIP 2
0.750
ConvNeXt-B
0.767
Table 1 : Disagreement prediction across anchor-model choices.
Table 2 : Disagreement prediction across comparison-model compositions and comparison model set sizes.
Aircraft
Caltech101
Cars
DTD
EuroSAT
Food101
Pets
SUN397
UCF101
Pearson r↑
0.661
0.656
0.425
0.587
0.405
0.585
0.499
0.597
0.637
Table 3 : Cross-dataset disagreement prediction. Pearson correlations between predicted and measured NAk on nine external datasets using the fixed ImageNet-1K-trained predictor. Dataset sources: Aircraft ( Maji et al., 2013 ) , Caltech101 ( Fei-Fei et al., 2004 ) , Cars ( Krause et al., 2013 ) , DTD ( Cimpoi et al., 2014 ) , EuroSAT ( Helber et al., 2019 ) , Food101 ( Bossard et al., 2014 ) , Pets ( Parkhi et al., 2012 ) , SUN397 ( Xiao et al., 2010 ) , and UCF101 ( Soomro et al., 2012 ) .
ImageNet-V2
ImageNet-R
ImageNet-C: Gaussian
Matched
Threshold 0.7
Top Images
S1
S2
S3
S4
S5
Pearson r↑
0.710
0.720
0.717
0.512
0.726
0.732
0.724
0.650
0.394
Table 4 : Disagreement prediction under ImageNet distribution shifts. Pearson correlations between predicted and measured NAk on ImageNet-V2 ( Recht et al., 2019 ) , ImageNet-R ( Hendrycks et al., 2021 ) , and Gaussian-corrupted ImageNet-C ( Hendrycks and Dietterich, 2019 ) . The labels S1, S2, S3, S4, and S5 denote corruption severity levels 1, 2, 3, 4, and 5, respectively.
Figure 4 : Prediction drift under Gaussian-noise corruption. From left to right, corruption severity increases; each example shows a query and its three nearest neighbors retrieved by CLIP, DINOv2, and ViT-21K.
Figure 5 : Leave-anchor-out diagnosis on ImageNet-1K. Both panels use only CLIP as input; the target includes CLIP on the left and excludes it on the right.
Model
2-layer MLP
Transformer Encoder
Residual FFN
CLIP
0.544
0.712
0.737
DINOv2
0.523
0.724
0.746
ViT-21K
0.490
0.696
0.708
Table 5 : Predictor architecture comparison measured by Pearson correlation. All predictors use the same single-representation input and measured consistency supervision.
Method
Time ↓ [-1pt] (ms/sample)
Peak memory ↓ [-1pt] (MiB)
Storage ↓ [-1pt] (MiB)
Measured (PyTorch)
49.54
20,643
10,009
Measured (FAISS)
20.54
29,729
10,009
Predicted
5.34
9
5
Table 6 : Post-embedding cost on 100K ImageNet-1K samples. Lower is better.
Aircraft
Caltech101
Cars
DTD
EuroSAT
Food101
Pets
SUN397
UCF-101
Average
Measured
0.24
0.20
0.09
0.16
0.09
0.32
0.16
0.20
0.18
0.18
Predicted [-1pt] +0.00
0.26 [-1pt] +0.02
0.23 [-1pt] +0.03
0.10 [-1pt] +0.01
0.17 [-1pt] +0.01
0.08 [-1pt] -0.01
0.40 [-1pt] +0.08
0.17 [-1pt] +0.01
0.23 [-1pt] +0.03
0.23 [-1pt] +0.05
0.21 [-1pt] +0.03
Table 7 : Predicted representational disagreement tracks downstream classification risk. We report sample-wise Kendall τb correlations between measured or predicted disagreement and multiclass Brier risk across nine benchmarks. Red values indicate increases over directly measured disagreement; green values indicate decreases.
Figure 6 : Reference set scaling for disagreement prediction. Pearson correlation retained relative to using the full ImageNet-1K training set as the reference set; all predictors use the same 50,000 training samples and are evaluated against the same full-reference NAk target.
Appendix figures & tables6 assets
Supplementary material from the paper’s appendix.
Appendix
Model
Model source
Representation
Dimension
CLIP ViT-B/32
openai/clip-vit-base-patch32
Projected image features
512
DINOv2 ViT-B/14
facebook/dinov2-base
Pooled output; CLS if unavailable
768
ViT-B/16
google/vit-base-patch16-224-in21k
Pooled output; CLS if unavailable
768
Swin-B
microsoft/swin-base-patch4-window7-224-in22k
Pooled output
1024
SigLIP 2 ViT-B/16
google/siglip2-base-patch16-224
Projected image features
768
ConvNeXt-B
timm/convnext_base.fb_in22k
Pooled pre-logit features
1024
Appendix
Table 8 : Foundation models and extracted representations.
Figure 7 : High (top) and low (bottom) cross-model neighborhood agreement. Rows correspond to CLIP, DINOv2, and ViT-21K.
Anchor model
Comparison models
Input dimension
CLIP
DINOv2, ViT-21K
512
DINOv2
CLIP, ViT-21K
768
ViT-21K
CLIP, DINOv2
768
Appendix
Table 9 : Configurations for prediction from a single representation.
Configuration
Comparison models
1
Swin, SigLIP 2, ConvNeXt
2
Swin, SigLIP 2, ViT-21K
3
CLIP, SigLIP 2, ConvNeXt
Appendix
Table 10 : Comparison-model compositions with DINOv2 fixed as the anchor.
Anchor model
Comparison models
CLIP
Swin, SigLIP 2, ConvNeXt
DINOv2
Swin, SigLIP 2, ConvNeXt
ViT-21K
Swin, SigLIP 2, ConvNeXt
Swin
DINOv2, ViT-21K, CLIP
SigLIP 2
DINOv2, ViT-21K, CLIP
ConvNeXt
DINOv2, ViT-21K, CLIP
Appendix
Table 11 : Configurations for evaluating anchor-model choice.
Number of models
Comparison set
2
Swin, SigLIP 2
3
Swin, SigLIP 2, ConvNeXt
4
Swin, SigLIP 2, ConvNeXt, DINOv2
5
Swin, SigLIP 2, ConvNeXt, DINOv2, ViT-21K
Appendix
Table 12 : Nested comparison sets with CLIP fixed as the anchor.
The existence of multiple, equally accurate models for a given predictive task leads to predictive multiplicity, where a Rashomon set of models achieve similar accuracy but diverge in their individual predictions. This inconsistency undermines trust in high-stakes applications where we want consistent predictions. We propose three approaches to reduce inconsistency among predictions for the members of the Rashomon set. The first approach is outlier correction. An outlier has a label that none of the good models are capable of predicting correctly. Outliers can cause the Rashomon set to have high variance predictions in a local area, so fixing them can lower variance. Our second approach is local patching. In a local region around a test point, models may disagree with each other because some of them are biased. We can detect and fix such biases using a validation set, which also reduces multiplicity. Our third approach is pairwise reconciliation, where we find pairs of models that disagree on a region around the test point. We modify predictions that disagree, making them less biased. These three approaches can be used together or separately, and they each have distinct advantages. The reconciled predictions can then be distilled into a single interpretable model for real-world deployment. In experiments across multiple datasets, our methods reduce disagreement metrics while maintaining competitive accuracy.
Parian Haghighat, Hadis Anahideh, Cynthia Rudin
University of Illinois Chicago, Chicago, IL, USA · Duke University, Durham, NC, USA
Human-model alignment is critical for trustworthy AI-assisted decision-making systems. Yet, most work evaluates model predictions against single ground-truth labels, overlooking that humans themselves often disagree on labels, a signal of genuine ambiguity. We investigate whether models struggle on the same instances that humans find difficult. We measure this on two vision datasets (FER+ and CIFAR-10H) where multiple human annotations per image capture human disagreement patterns. We evaluate eight pretrained models across three architectures (ResNet, EfficientNet, MobileNetV3) in two parts: first, whether model uncertainty (softmax confidence, entropy) correlates with human disagreement, and second, whether predictive multiplicity measures (inter-model disagreement, Jensen-Shannon divergence) do. We find that it does not: alignment is weak in both dimensions. At the discrete label level, 50.4% of CIFAR-10H images and 33.5% of FER+ images receive multiple valid classifications from humans, while the models converge on only one. These instances represent a critical failure case where humans perceive ambiguity and would request expert review, yet models decide confidently. At the continuous score level, single-model uncertainty correlates weakly with human disagreement (ρ=0.24−−0.55), and predictive multiplicity provides only modest improvement. Widely-used uncertainty quantification methods do not reliably identify instances humans find ambiguous. Model uncertainty should not be treated as a trustworthy signal by default for decision-making in high-stakes scenarios.
Manya Singh, Arjun Pakrashi
School of Computer Science, University College Dublin, Ireland
Large language models trained under diverse objectives and architectures have been shown to develop increasingly similar internal representations, an observation formalized as the Platonic Representation Hypothesis. Whether this representational convergence extends to the reasoning processes that operate over shared representations remains untested. We evaluate representational similarity across 16 language models from 8 families (1.5B to 72B parameters) on 800 reasoning problems spanning mathematics, science, commonsense, and truthfulness, stratifying by problem difficulty, computational stage, and causal relevance. Our analysis reveals three dissociations: a difficulty inversion, where models converge more on problems they collectively fail (Centered Kernel Alignment [CKA] = 0.897) than on those they solve (CKA = 0.830); a generation gap, where pre-decision representations align (CKA = 0.875) while post-decision representations diverge (CKA = 0.274); and epiphenomenal correctness, where shared information is decodable across models (66% transfer accuracy) but exerts minimal causal influence on predictions (1.5% to 5.5% flip rate across ablation protocols). These results indicate that representational convergence in language models reflects shared input processing constraints rather than shared reasoning strategies, with direct implications for ensemble design, interpretability transfer, and evaluations of model similarity. Code is available at https://github.com/Usama1002/convergence-without-understanding.
Muhammad Usama, Dong Eui Chang
Control Laboratory, School of Electrical Engineering Korea Advanced Institute of Science and Technology (KAIST) Daejeon 34141, Republic of Korea · Control Laboratory, School of Electrical Engineering Korea Advanced Institute of Science and Technology (KAIST)2026 Daejeon 34141, Republic of Korea