cs.CVSep 26, 2026

Multimodal LLMs Outperform Pathology Foundation Models in Cross-Domain Histological Similarity

Authors: Yishu Zhang, Yun Li, Daiwei Zhang

Organizations: University of North Carolina at Chapel Hill

Abstract

State-of-the-art pathology foundation models, trained on millions of histology tiles, can fail to preserve tissue similarity when comparisons cross slide or institution boundaries. We show that general-purpose multimodal LLMs, without being trained as pathology foundation models, consistently outperform these specialized models in cross-domain histological similarity judgments. Using a relative similarity framework that we release as the MOSAIC (Model Similarity Assessment across Institutions and Cohorts) benchmark, we evaluate 17 models across 6 datasets and find that pathology encoders often rank same-institution, different-disease tiles as more similar than same-disease, different-institution tiles, a clinically dangerous failure mode invisible to standard within-domain evaluations. LLMs appear less susceptible to this failure, likely because they perform semantic visual comparison of morphology and tissue architecture rather than relying on shortcut features tied to acquisition context. Scaling training data does not resolve the problem for pathology encoders, implicating the learning objective rather than data coverage. Our results expose a fundamental robustness gap in current pathology foundation models and establish multimodal LLMs as a viable alternative for cross-institutional retrieval, dataset harmonization, and multi-site quality control. Code and data will be released upon acceptance.

Figures & tables

Appendix figures & tables24 assets

Supplementary material from the paper’s appendix.

Appendix

Explore similar work

CardsList
  1. ALICE: Learning a General-Purpose Pathology Foundation Model from Vision, Vision-Language, and Slide-Level Experts

    Jul 10, 2026Jiawen Li, Tian Guan, Huijuan Shi +5Pathology Foundation ModelsProbe-Logit Distillation

  2. How Seemingly Inconsequential Design Choices Dictate Performance of LLMs in Pathology

    Jun 10, 2026Kian R. Weihrauch, Thomas A. Buckley, William Lotter +1Pathology Foundation ModelsModel Size

  3. The Good, the Bad, and the Brittle: Benchmarking Robustness and Generalisation of Histopathology Foundation Models

    Jul 5, 2026Dhyey Yajnik, Amina Asif, Fayyaz MinhasPathology Foundation ModelsComputational Pathology