cs.CVJul 5, 2026

The Good, the Bad, and the Brittle: Benchmarking Robustness and Generalisation of Histopathology Foundation Models

Authors: Dhyey YajnikAmina AsifFayyaz Minhas

Organizations: Predictive Systems in Biomedicine Lab, Tissue Image Analytics Centre, Department of Computer Science, University of Warwick, United Kingdom

Abstract

How robust and generalisable are pathology foundation models and have their scaling limites been reached? We benchmarked twelve pathology foundation models (PFMs) and ResNet baselines using our Robustness Evaluation and Enhancement Toolbox (REET) across eleven clinically realistic perturbations and a dissimilarity-driven Non-Redundant K-fold validation (NR-Kfold) protocol. We introduce a Perturbation Performance Index (PPI) to summarise accuracy trends under controlled perturbation sweeps and analyse robustness scaling with parameter count. We show that PFMs consistently outperform CNNs in both robustness and domain generalisation, yet model scaling shows diminishing returns: mid-sized models such (UNI2/Virchow-2 etc.) achieve comparable or greater resilience than larger systems. NR-Kfold analysis further reveals systematic accuracy loss and increased variability when training-test similarity is broken, underscoring the need for explicit distribution-shift evaluation. These findings suggest that the next generation of pathology foundation models must prioritise data quality, multimodality information and domain alignment over parameter count to achieve genuine clinical reliability.

Explore similar work

CardsList
  1. Robustifying pathology foundation models via fine-tuning

    Jul 24, 2026Alexandre Filiot, Oskar Thaeter, Benoit Schmauch +1Pathology Foundation Models