3D Masked Autoencoders are Robust Learners of Volumetric and Multimodal Cellular Representations for Microscopy
Organizations: Institute of AI for Health & Helmholtz AI, Computational Health Center, Helmholtz Munich – German Research Center for Environmental Health, Neuherberg, Germany · Munich Center for Machine Learning (MCML), Munich, Germany · Department of Medicine III, Ludwig-Maximilian-University Hospital, Munich, Germany · Department of Physics, Ludwig-Maximilian-University, Munich, Germany · German Cancer Consortium (DKTK), partner site Munich, Germany
Abstract
Self-supervised learning in fluorescence microscopy often relies on 2D projections, despite the inherently three-dimensional nature of cells. We present a systematic comparison of 2D and 3D masked autoencoders (MAE-2D vs. MAE-3D) on volumetric microscopy data. Under matched architectures and training protocols, MAE-3D consistently outperforms 2D max-projection and slice-based variants on downstream single-cell tasks. We further align visual representations with a pretrained protein language model (ESM2) and show that cross-modal supervision yields larger gains for volumetric models. Channel cross-attention and frequency-domain regularization are critical for leveraging 3D spatial context. On protein--protein interaction prediction, our best model achieves a ROC--AUC of 0.86, while on protein localization it reaches an AUC of 0.95 and an F1 of 0.74, demonstrating competitive performance on both tasks. Overall, our findings highlight the potential of volumetric modeling and multimodal alignment for representation learning in single-cell microscopy.