cs.LGJun 22, 2026

3D Masked Autoencoders are Robust Learners of Volumetric and Multimodal Cellular Representations for Microscopy

Authors: Amirhossein KardoostLion GleiterTingying PengCarsten Marr

Organizations: Institute of AI for Health & Helmholtz AI, Computational Health Center, Helmholtz Munich – German Research Center for Environmental Health, Neuherberg, Germany · Munich Center for Machine Learning (MCML), Munich, Germany · Department of Medicine III, Ludwig-Maximilian-University Hospital, Munich, Germany · Department of Physics, Ludwig-Maximilian-University, Munich, Germany · German Cancer Consortium (DKTK), partner site Munich, Germany

Abstract

Self-supervised learning in fluorescence microscopy often relies on 2D projections, despite the inherently three-dimensional nature of cells. We present a systematic comparison of 2D and 3D masked autoencoders (MAE-2D vs. MAE-3D) on volumetric microscopy data. Under matched architectures and training protocols, MAE-3D consistently outperforms 2D max-projection and slice-based variants on downstream single-cell tasks. We further align visual representations with a pretrained protein language model (ESM2) and show that cross-modal supervision yields larger gains for volumetric models. Channel cross-attention and frequency-domain regularization are critical for leveraging 3D spatial context. On protein--protein interaction prediction, our best model achieves a ROC--AUC of 0.86, while on protein localization it reaches an AUCmicro_{\text{micro}} of 0.95 and an F1micro_{\text{micro}} of 0.74, demonstrating competitive performance on both tasks. Overall, our findings highlight the potential of volumetric modeling and multimodal alignment for representation learning in single-cell microscopy.

Explore similar work

CardsList