cs.LGDec 29, 2025

Stochastic Siamese MAE Pretraining for Longitudinal Medical Images

Authors: Taha Emre, Arunava Chakravarty, Thomas Pinetz, Dmitrii Lachinov, Martin J. Menten, Hendrik Scholl, Sobha Sivaprasad, Daniel Rueckert, +4 more

Organizations: Institute of Artificial Intelligence, Center for Medical Data Science, Medical University of Vienna, Austria · Department of Ophthalmology and Optometry, Medical University of Vienna, Austria · BioMedIA, Department of Computing, Imperial College London, London, United Kingdom · Chair for AI in Healthcare and Medicine, Technical University of Munich, Munich, Germany · Department of Clinical Pharmacology, Medical University of Vienna, Vienna, Austria · Pallas Kliniken AG, Pallas Klinik Zürich, Zürich, Switzerland · European Vision Institute, Basel, Basel-Stadt, Switzerland · Moorfields National Institute for Health and Care Biomedical Research Centre, Moorfields Eye Hospital, London, United Kingdom · Institute of Ophthalmology, University College London, London, United Kingdom · Faculty of Medicine, University of Southampton, Southampton, Hampshire, United Kingdom · Ophthalmic Image Analysis Group (OPTIMA), Medical University of Vienna, Austria

Abstract

Temporally aware image representations are crucial for capturing disease progression in 3D volumes of longitudinal medical datasets. However, recent state-of-the-art self-supervised learning approaches like Masked Autoencoding (MAE), despite their strong representation learning capabilities, lack temporal awareness. In this paper, we propose STAMP (Stochastic Temporal Autoencoder with Masked Pretraining), a Siamese MAE framework that encodes temporal information through a stochastic process by conditioning on the time difference between the 2 input volumes. Unlike deterministic Siamese approaches, which compare scans from different time points but fail to account for the inherent uncertainty in disease evolution, STAMP learns temporal dynamics stochastically by reframing the MAE reconstruction loss as a conditional variational inference objective. We evaluated STAMP on two OCT and one MRI datasets with multiple visits per patient. STAMP pretrained ViT models outperformed both existing temporal MAE methods and foundation models on different late stage Age-Related Macular Degeneration and Alzheimer's Disease progression prediction which require models to learn the underlying non-deterministic temporal dynamics of the diseases.

Figures & tables

Explore similar work

CardsList
  1. AD-DAE: Alzheimer's Disease Progression Modeling with Unpaired Longitudinal MRI using Diffusion Auto-Encoders

    Nov 8, 2025Ayantika Das, Arunima Sarkar, Keerthi Ram +1AlzheimerReadmission Prediction

  2. CLIMB: Controllable Longitudinal Brain Image Generation using Mamba-based Latent Diffusion Model and Gaussian-aligned Autoencoder

    Apr 17, 2026Duy-Phuong Dao, Muhammad Taqiyuddin, Jahae Kim +4Identity-Conditioned Latent Diffusion ModelsLatent Diffusion Model

  3. Self-supervised Pre-training Helps Retinal Disease Progression Modelling Most When Data Is Scarce

    Sep 14, 2026Ifeoma Veronica Nwabufo, Julius Gervelmeyer, Sarah Müller +1Self-Supervised LearningLongitudinal Imaging