cs.CVAug 1, 2026

Representation Transfer of Foundation Models for Ultra-Widefield Retinal Imaging

Authors: Mingya Alexa GongDa MaLovre Antonio BudimirIvana MatovinovicSven LoncaricMyeong Jin JuYukun ZhouSiegfried K. Wagner+2 more

Organizations: Institute of Ophthalmology, University College London, London, United Kingdom · Wake Forest University School of Medicine, Winston-Salem, NC, USA · Virginia Tech-Wake Forest University School of Biomedical Engineering and Sciences, Blacksburg, VA, USA · University of Zagreb Faculty of Electrical Engineering and Computing, Zagreb, Croatia. · Department of Ophthalmology and Visual Sciences, University of British Columbia, Vancouver, BC, Canada · School of Biomedical Engineering, University of British Columbia, Vancouver, BC, Canada · NIHR Biomedical Research Centre, Moorfields Eye Hospital NHS Foundation Trust, London, United Kingdom · Department of Medical Physics and Biomedical Engineering, University College London, London, United Kingdom

Abstract

Despite the widespread adoption of foundation models as feature extractors for medical imaging, relatively little is understood about how different pretraining strategies influence the transferability of learned representations to weakly supervised ophthalmic imaging tasks. We investigate this question in ultra-widefield (UWF) retinal imaging by evaluating foundation model representations within a patch-based multiple instance learning (MIL) framework for disease classification on UWF images. We compare Vision Transformer encoders pretrained with supervised, Masked Autoencoder (MAE), and self-distillation objectives, while keeping the downstream aggregation architecture unchanged. Within a controlled comparison of ViT-B encoders pretrained on ImageNet-1k, the choice of pretraining objective substantially influenced frozen representation transfer, with supervised and self-distillation-based models outperforming MAE. A contemporary DINOv3 model pretrained at a larger scale achieved the strongest overall performance, with a quadratic weighted kappa of 0.863 for five-class diabetic retinopathy grading, comparable with DINOv1. Attention analysis further revealed distinct patch-aggregation behaviours associated with the different pretrained representations, while partial fine-tuning substantially reduced the performance gap for MAE. These findings suggest that pretraining strategy influences both representation transferability and the subsequent aggregation of patch-level evidence within MIL, resulting in differences in downstream classification performance.

Explore similar work

CardsList
  1. Towards Interpretable Foundation Models for Retinal Fundus Images

    Mar 19, 2026Samuel Ofosu Mensah, Camila Roa, Kerol Djoumessi +1Fundus ImagesRetinal Imaging