Multimodal LLMs Outperform Pathology Foundation Models in Cross-Domain Histological Similarity
Organizations: University of North Carolina at Chapel Hill
Abstract
State-of-the-art pathology foundation models, trained on millions of histology tiles, can fail to preserve tissue similarity when comparisons cross slide or institution boundaries. We show that general-purpose multimodal LLMs, without being trained as pathology foundation models, consistently outperform these specialized models in cross-domain histological similarity judgments. Using a relative similarity framework that we release as the MOSAIC (Model Similarity Assessment across Institutions and Cohorts) benchmark, we evaluate 17 models across 6 datasets and find that pathology encoders often rank same-institution, different-disease tiles as more similar than same-disease, different-institution tiles, a clinically dangerous failure mode invisible to standard within-domain evaluations. LLMs appear less susceptible to this failure, likely because they perform semantic visual comparison of morphology and tissue architecture rather than relying on shortcut features tied to acquisition context. Scaling training data does not resolve the problem for pathology encoders, implicating the learning objective rather than data coverage. Our results expose a fundamental robustness gap in current pathology foundation models and establish multimodal LLMs as a viable alternative for cross-institutional retrieval, dataset harmonization, and multi-site quality control. Code and data will be released upon acceptance.
Figures & tables
| HER2ST | CCTGS | TIGER | ||||||||||||||||
| Method | Adi | Bre | Can | Con | Imm | All | Adi | Lam | Lym | Mus | Nor | Tum | All | Hea | Nec | Str | Tum | All |
| Pathology foundation models | ||||||||||||||||||
| UNI2H | 0.61 | 0.57 | 0.38 | 0.38 | 0.36 | 0.46 | 0.60 | 0.58 | 0.56 | 0.62 | 0.55 | 0.62 | 0.59 | 0.39 | 0.21 | 0.09 | 0.13 | 0.19 |
| H-optimus-0 | 0.59 | 0.53 | 0.36 | 0.40 | 0.32 | 0.44 | 0.59 | 0.56 | 0.57 | 0.59 | 0.51 | 0.51 | 0.55 | 0.40 | 0.20 | 0.17 | 0.17 | 0.22 |
| Midnight | 0.60 | 0.56 | 0.41 | 0.37 | 0.40 | 0.46 | 0.76 | 0.70 | 0.69 | 0.76 | 0.75 | 0.68 | 0.72 | 0.50 | 0.12 | 0.11 | 0.16 | 0.20 |
| CONCH v1.5 | 0.43 | 0.52 | 0.51 | 0.41 | 0.38 | 0.45 | 0.47 | 0.59 | 0.58 | 0.66 | 0.53 | 0.55 | 0.56 | 0.32 | 0.19 | 0.18 | 0.20 | 0.21 |
| Method | CAMELYON16 | TCGA | BEETLE |
|---|---|---|---|
| Pathology foundation models | |||
| UNI2H | 0.61 | 0.31 | 0.56 |
| H-optimus-0 | 0.40 | 0.18 | 0.45 |
| Midnight | 0.64 | 0.36 | 0.48 |
| CONCH v1.5 | 0.57 | 0.37 | 0.47 |
| Virchow2 | 0.63 | 0.44 | 0.57 |
| HER2ST (cross-slide) | CAMELYON16 (cross-inst.) | |||
|---|---|---|---|---|
| Prompt | Gemini 3 Flash | GPT-5 Nano | Gemini 3 Flash | GPT-5 Nano |
| Biology-focused (default) | 0.549 | 0.570 | 0.705 | 0.590 |
| Minimal | 0.510 | 0.560 | 0.668 | 0.555 |
Appendix figures & tables24 assets
Supplementary material from the paper’s appendix.
Appendix
| Method | Adipose | Breast gl. | Cancer | Connective | Immune inf. | All |
|---|---|---|---|---|---|---|
| Pathology foundation models | ||||||
| UNI2H | 0.775 | 0.759 | 0.664 | 0.757 | 0.712 | 0.728 |
| H-optimus-0 | 0.709 | 0.691 | 0.682 | 0.748 | 0.756 | 0.719 |
| Midnight | 0.706 | 0.732 | 0.670 | 0.695 | 0.700 | 0.697 |
| CONCH v1.5 | 0.706 | 0.736 | 0.691 | 0.689 | 0.724 | 0.705 |
| Virchow2 | 0.741 | 0.718 | 0.755 | 0.739 | 0.709 | 0.735 |
| Method | Radboud normal | Radboud tumor | Utrecht normal | Utrecht tumor | All |
|---|---|---|---|---|---|
| Pathology foundation models | |||||
| UNI2H | 0.340 | 0.770 | 0.440 | 0.900 | 0.613 |
| H-optimus-0 | 0.390 | 0.310 | 0.490 | 0.400 | 0.398 |
| Midnight | 0.610 | 0.670 | 0.550 | 0.730 | 0.640 |
| CONCH v1.5 | 0.460 | 0.630 | 0.410 | 0.760 | 0.565 |
| Virchow2 | 0.500 | 0.660 | 0.630 | 0.710 | 0.625 |
| Method | TSS-22 normal | TSS-22 tumor | TSS-56 normal | TSS-56 tumor | TSS-66 normal | TSS-66 tumor | All |
|---|---|---|---|---|---|---|---|
| Pathology foundation models | |||||||
| UNI2H | 0.190 | 0.540 | 0.220 | 0.450 | 0.090 | 0.380 | 0.312 |
| H-optimus-0 | 0.020 | 0.040 | 0.200 | 0.300 | 0.200 | 0.300 | 0.177 |
| Midnight | 0.110 | 0.080 | 0.470 | 0.590 | 0.300 | 0.590 | 0.357 |
| CONCH v1.5 | 0.300 | 0.410 | 0.300 | 0.480 | 0.300 | 0.400 | 0.365 |
| Virchow2 | 0.330 | 0.470 | 0.300 | 0.530 | 0.430 | 0.580 | 0.440 |
| Method | NKI inv. | NKI nec. | NKI non-inv. | RUMC inv. | RUMC nec. | RUMC non-inv. | SCH inv. | SCH nec. | SCH non-inv. | TCGA inv. | TCGA nec. | TCGA non-inv. | All |
|---|---|---|---|---|---|---|---|---|---|---|---|---|---|
| Pathology foundation models | |||||||||||||
| UNI2H | 0.650 | 0.750 | 0.800 | 0.620 | 0.670 | 0.670 | 0.300 | 0.310 | 0.320 | 0.400 | 0.680 | 0.520 | 0.557 |
| H-optimus-0 | 0.600 | 0.570 | 0.430 | 0.800 | 0.410 | 0.490 | 0.410 | 0.310 | 0.390 | 0.500 | 0.170 | 0.280 | 0.447 |
| Midnight | 0.350 | 0.740 | 0.860 | 0.120 | 0.630 | 0.700 | 0.140 | 0.420 | 0.590 | 0.190 | 0.290 | 0.700 | 0.477 |
| CONCH v1.5 | 0.450 | 0.470 | 0.450 | 0.700 | 0.710 | 0.590 | 0.390 | 0.370 | 0.360 | 0.360 | 0.310 | 0.430 | 0.466 |
| Virchow2 | 0.390 | 0.810 | 0.740 | 0.660 | 0.520 | 0.850 | 0.600 | 0.170 | 0.480 | 0.420 | 0.420 | 0.720 | 0.565 |
| Method | Radboud normal | Radboud tumor | Utrecht normal | Utrecht tumor | All |
|---|---|---|---|---|---|
| Pathology foundation models | |||||
| UNI2H | 0.610 | 0.970 | 0.530 | 0.940 | 0.763 |
| H-optimus-0 | 0.570 | 0.890 | 0.680 | 0.890 | 0.758 |
| Midnight | 0.730 | 0.970 | 0.680 | 0.930 | 0.828 |
| CONCH v1.5 | 0.610 | 0.950 | 0.560 | 0.910 | 0.758 |
| Virchow2 | 0.710 | 0.960 | 0.790 | 0.950 | 0.853 |
| Method | TSS-22 normal | TSS-22 tumor | TSS-56 normal | TSS-56 tumor | TSS-66 normal | TSS-66 tumor | All |
|---|---|---|---|---|---|---|---|
| Pathology foundation models | |||||||
| UNI2H | 0.730 | 0.760 | 0.430 | 0.690 | 0.490 | 0.830 | 0.655 |
| H-optimus-0 | 0.720 | 0.820 | 0.450 | 0.730 | 0.460 | 0.920 | 0.683 |
| Midnight | 0.760 | 0.750 | 0.520 | 0.650 | 0.570 | 0.890 | 0.690 |
| CONCH v1.5 | 0.710 | 0.600 | 0.450 | 0.690 | 0.470 | 0.870 | 0.632 |
| Virchow2 | 0.780 | 0.720 | 0.500 | 0.720 | 0.540 | 0.840 | 0.683 |
| Method | NKI inv. | NKI nec. | NKI non-inv. | RUMC inv. | RUMC nec. | RUMC non-inv. | SCH inv. | SCH nec. | SCH non-inv. | TCGA inv. | TCGA nec. | TCGA non-inv. | All |
|---|---|---|---|---|---|---|---|---|---|---|---|---|---|
| Pathology foundation models | |||||||||||||
| UNI2H | 0.930 | 0.940 | 0.980 | 0.770 | 1.000 | 0.945 | 0.945 | 0.905 | 0.740 | 0.785 | 1.000 | 0.780 | 0.893 |
| H-optimus-0 | 0.885 | 0.950 | 0.965 | 0.845 | 1.000 | 0.910 | 0.985 | 0.795 | 0.715 | 0.815 | 1.000 | 0.795 | 0.888 |
| Midnight | 0.940 | 0.965 | 0.950 | 0.870 | 1.000 | 0.945 | 0.930 | 0.905 | 0.770 | 0.905 | 0.995 | 0.755 | 0.910 |
| CONCH v1.5 | 0.840 | 0.960 | 0.805 | 0.870 | 0.995 | 0.905 | 0.970 | 0.825 | 0.640 | 0.765 | 0.975 | 0.755 | 0.859 |
| Virchow2 | 0.945 | 0.955 | 0.970 | 0.835 | 1.000 | 0.975 | 0.980 | 0.870 | 0.800 | 0.790 | 0.975 | 0.765 | 0.905 |
| Model | Training size | Source | URL | Quote |
| Pathology foundation models | ||||
| UNI2H | 350K WSIs | HuggingFace model card | https://huggingface.co/MahmoodLab/UNI2-h | “Over 200 million image tiles sampled from over 350k diverse H&E and IHC slides sourced from Mass General Brigham.” |
| H-optimus-0 | 500K WSIs | HuggingFace model card | https://huggingface.co/bioptimus/H-optimus-0 | “The model is a 1.1B parameter vision transformer trained on a proprietary collection of more than 500,000 H&E stained whole slide histology images.” |
| Midnight | 12K WSIs | Tolkach et al. (2025) | https://arxiv.org/abs/2504.05186 | “We trained our first FM on the 12k TCGA WSIs alone.” |
| CONCH v1.5 | 1.26M pairs | Ding et al. (2024) | https://arxiv.org/abs/2411.19666 | “CONCHv1.5, an extended version of CONCH, which was trained with 1.26 million image-caption pairs using the CoCa training objective.” |
| Virchow2 | 3.1M WSIs | Zimmermann et al. (2024) | https://arxiv.org/abs/2408.00738 | “each trained with 3.1 million histopathology whole slide images” |
| Model | Price per 1M input tokens (USD) |
|---|---|
| Gemini 3 Flash | $0.50 |
| Gemini 3.1 Flash Lite | $0.25 |
| Qwen3.5 | $0.39 |
| GLM-4.6V | $0.30 |
| Gemma 3 | $0.08 |
| GPT-5 Nano | $0.05 |
| HER2ST (cross-slide) | CAMELYON16 (cross-inst.) | |||||
| Method | Acc. | 95% CI | Acc. | 95% CI | ||
| Pathology foundation models | ||||||
| UNI2H | 0.459 | [0.437, 0.480] | 2000 | 0.613 | [0.565, 0.660] | 400 |
| H-optimus-0 | 0.439 | [0.416, 0.461] | 2000 | 0.398 | [0.350, 0.445] | 400 |
| Midnight | 0.464 | [0.442, 0.485] | 2000 | 0.640 | [0.593, 0.685] | 400 |
| CONCH v1.5 | 0.450 | [0.428, 0.472] | 2000 | 0.565 | [0.517, 0.613] | 400 |
| Experiment | Comparison | Acc. A | Acc. B | -value | |
|---|---|---|---|---|---|
| HER2ST (cross-slide) | GPT-5 Nano vs Midnight | 0.570 | 0.464 | 58.1 | 0.0001 |
| GPT-5 Nano vs DINOv2 | 0.570 | 0.549 | 2.0 | 0.153 | |
| CAMELYON16 (cross-inst.) | Gem. 3.1 F. Lite vs Midnight | 0.730 | 0.640 | 11.8 | 0.001 |
| Gem. 3.1 F. Lite vs Gemini Emb. 2 | 0.730 | 0.610 | 26.9 | 0.0001 |
| Dataset | Abbreviation | Full name |
| HER2ST | Adi | Adipose |
| Bre | Breast glandular | |
| Can | Cancer | |
| Con | Connective | |
| Imm | Immune infiltrate | |
| CCTGS | Adi | Adipose |
| Model | License |
|---|---|
| Pathology foundation models | |
| UNI2H | CC-BY-NC-ND 4.0 |
| H-optimus-0 | Apache 2.0 |
| Midnight | MIT |
| CONCH v1.5 | CC-BY-NC-ND 4.0 |
| Virchow2 | CC-BY-NC-ND 4.0 |