Benchmarking the Robustness of Foundation Models for Mammography under Domain Shift
Authors: Giang Nguyen, Raghav Mehta, Emma A. M. Stanley, Tian Xia, Thi Hao Nguyen, Hieu Pham, Ben Glocker
Organizations: College of Engineering and Computer Science, VinUniversity, Hanoi, Vietnam · Imperial College London, London, UK · Radiology Department, Vietnam National Cancer Hospital, Hanoi, Vietnam · VinUni-Illinois Smart Health Center, VinUniversity, Hanoi, Vietnam · The Computer Vision and Medical AI Lab, VinUniversity, Hanoi, Vietnam
Foundation models are increasingly used as image feature extractors for mammography, but their robustness under external domain shift remains unclear. We benchmark 15 foundation-model backbones across breast density, BI-RADS severity, and cancer status using a unified frozen-backbone linear-probe protocol, training on 3 source datasets and evaluating on 12 task-compatible out-of-distribution (OOD) datasets after label harmonization. Mammography-specific vision-language models (Mammo-FM and MaMA) provide the strongest mean OOD performance, but robustness is not explained by mammography exposure alone. DINOv3 remains a competitive vision-only baseline, and mammography-adapted pretraining does not consistently improve generalization. Dataset-level analysis further shows that even leading models show heterogeneous performance across datasets. Feature-space inspection reveals that useful representations can preserve clinical signal while retaining dataset and acquisition structure. These findings highlight dataset-level OOD evaluation as a central criterion for assessing mammography representations. Our code is publicly available: https://github.com/biomedia-mira/mammo-ood.