stat.MLOct 1, 2026

The hidden advantage of mask resampling: a theory of masked autoencoders

Authors: Jorge Medina Moreira, Lorenzo Bardone, Lenka Zdeborová

Organizations: Statistical Physics of Computation Laboratory, École Polytechnique Fédérale de Lausanne (EPFL), CH-1015 Lausanne, Switzerland

Abstract

Why can masked prediction learn useful representations that unmasked reconstruction misses? We study this question in a high-dimensional model of a masked autoencoder (MAE) trained on data with shared latent structure and heterogeneous noise. We prove that masked linear reconstruction can recover the latent feature at linear sample complexity in regimes where unmasked linear reconstruction, equivalent to PCA, fails. The analysis also quantifies the statistical advantage of mask resampling, an established ingredient of masked pretraining. By introducing a fixed collection of KK masks per sample, we characterize its effect on feature recovery and downstream performance, identifying regimes where greater mask diversity lowers sample complexity. Guided by this prediction, we find that random cropping and flipping in standard image-training pipelines can obscure the advantage of mask resampling by renewing the prediction task even when the patch mask is fixed. Removing these transformations reveals a downstream advantage for dynamic over static masking in CNN autoencoders and vision transformers. A complementary BERT pilot finds benefits from greater mask diversity on downstream language tasks. Our results separate the benefit of the masked prediction objective from that of mask diversity, and show how a tractable theory can guide experiments that uncover advantages hidden by standard training practices.

Figures & tables

Appendix figures & tables17 assets

Supplementary material from the paper’s appendix.

Appendix

Explore similar work

CardsList
  1. Masked Swingers: Harnessing Data Augmentation to Advance Autoencoders for Self-Supervised Learning

    Sep 29, 2026Anthony Fuller, Scott C. Lowe, Daniel G. Kyrollos +3Self-Supervised LearningAutoencoder Architectures

  2. An Identifiability Theory of Masked Prediction: Mode Blindness and Mask Schedules

    Aug 2, 2026Yichao Cai, Javen Qinfeng ShiIdentifiabilityPredictability

  3. Mitigating Positional Leakage in 3D Masked Autoencoders for Robust Representation Learning

    Jun 30, 2026Xu Yan, Huiqun Wang, Chen Wang +23D Masked AutoencodersAutoencoder Architectures