Why can masked prediction learn useful representations that unmasked reconstruction misses? We study this question in a high-dimensional model of a masked autoencoder (MAE) trained on data with shared latent structure and heterogeneous noise. We prove that masked linear reconstruction can recover the latent feature at linear sample complexity in regimes where unmasked linear reconstruction, equivalent to PCA, fails. The analysis also quantifies the statistical advantage of mask resampling, an established ingredient of masked pretraining. By introducing a fixed collection of K masks per sample, we characterize its effect on feature recovery and downstream performance, identifying regimes where greater mask diversity lowers sample complexity. Guided by this prediction, we find that random cropping and flipping in standard image-training pipelines can obscure the advantage of mask resampling by renewing the prediction task even when the patch mask is fixed. Removing these transformations reveals a downstream advantage for dynamic over static masking in CNN autoencoders and vision transformers. A complementary BERT pilot finds benefits from greater mask diversity on downstream language tasks. Our results separate the benefit of the masked prediction objective from that of mask diversity, and show how a tractable theory can guide experiments that uncover advantages hidden by standard training practices.
Figures & tables
Figure 1: Isolating mask resampling reveals its contribution to representation learning. Downstream accuracy versus masking ratio ρ for dynamic masking (blue), static masking (yellow), and unmasked reconstruction (red). (a) Shallow autoencoder on spiked model ( Sec. 2 ); curves are theory and markers are simulations. (b) CNN autoencoder on CIFAR-10 (test top-1). (c) ViT-Base on ImageNet-100 (validation top-5). In (b,c), blue and yellow curves use no stochastic cropping or flipping. The purple reference includes cropping and flipping with dynamic mask resampling at ρ=0.625 (b) and ρ=0.75 (c) . At suitable masking ratios, resampling alone reaches comparable accuracy. Bars and shading show ±1 standard deviation (SD). See App. H for details.
Figure 2: Masking recovers a feature that PCA misses. Cosine similarity versus sample complexity α=n/d , for (a) homogeneous noise, γ=1 ; (b) bounded heterogeneous noise, γ=41+49B , B∼Beta(1,2) ; and (c) unbounded heterogeneous noise, γ=43+41eZ−1/2 , Z∼N(0,1) . Curves are theoretical predictions (linear masked autoencoders), markers are finite-dimensional simulations, and error bars show ±1 SD across seeds. Green curves give the matched Bayes-optimal reference; green diamonds show AMP results ( Secs. F.5 and H.14 ). Unmasked reconstruction is ordinary PCA. The homogeneous panel provides a control; the heterogeneous panels illustrate recovery by masking where the asymptotic PCA alignment vanishes. See App. H for details.
Figure 3: Replica predictions track learning across mask diversity. Cosine similarity (top) and downstream classification accuracy (bottom) versus sample complexity α=n/d under bounded heterogeneous noise (variance distribution of Fig. 2 (b)), for linear, ReLU, ELU, and tanh activations. Colors distinguish K=1 (static) ,2,4 and the dynamic masking limit. Solid curves are replica predictions and markers are finite-dimensional simulations. Green curves show the Bayes-optimal pretraining reference, evaluated with the same logistic probe for the lower row; green diamonds show AMP results ( Secs. F.5 and H.14 ). Error bars show ±1 SD. See App. H for experimental details.
Figure 4: Finite mask diversity interpolates between static and dynamic training. Downstream accuracy versus masking ratio for (a) the solvable model, with theoretical curves and simulation markers, and (b) a CNN autoencoder on CIFAR-10. Increasing K improves accuracy in the settings shown, and the curves approach the dynamic masking reference. The number of training samples is held fixed. Bars and shading show ±1 SD across seeds; see App. H for details.
Figure 5: Mask renewal and image augmentation. Linear-probe validation top-5 accuracy for ViT-Base on ImageNet-100. Blue/yellow indicate dynamic / static masking; circles/crosses denote pretraining without/with stochastic cropping and flipping, respectively. All four conditions use a common seed, data split, initialization, optimizer settings, and linear-probing protocol. Each point represents one run. The blue cross at ρ=0.75 is the purple horizontal line in Fig. 1 (c). See App. H for training details.
Masking
MNLI-m
QNLI
QQP
SST-2
AVG
K=1
72.9
82.6
85.9
87.6
62.7
K=10
76.3
84.4
87.0
88.1
65.8
Dynamic
76.8
85.5
87.2
89.4
66.4
Table 1: Development scores ( ×100 , ↑ ) for BERT-Medium under different masking procedures, on a subset of GLUE tasks. AVG is the unweighted mean of the task-specific scores over all eight evaluated tasks (accuracy; Spearman correlation for STS-B; MCC for MRPC, CoLA, and RTE), using MNLI-m for MNLI. Entries are means over three pretraining seeds, each averaged over two fine-tuning runs; standard deviations are in Table 5 . See Sec. H.13 for details.
Appendix figures & tables17 assets
Supplementary material from the paper’s appendix.
Appendix
Variance law
αcstat
αcdyn
τcdyn
vc(γmin)/vc(γmax)
Balanced
2.28272
0.52222
0.19380
1.24442
Rare high variance
2.28272
0.43843
0.26078
1.87089
Very clean component
2.28272
0.54817
0.17987
1.25380
Appendix
Table 2: Identical static thresholds, different dynamic adaptation. Regular linear replica thresholds for the laws in ( D.154 ) at β=1.5 , ρ=0.95 , and zero ridge. All three laws have unit mean and κ2=1.45 . The response ratio compares the lowest- and highest-variance coordinates using ( D.143 ). Entries are computed numerically from the replica equations.
Pretraining
Linear probe
Accuracy (%, ↑ )
Schedule
Epochs
Batch
Epochs
Batch
Top-1
Top-5
Main comparison
100
128
50
256
39.96
68.78
U-MAE
200
1024
50
256
47.56
76.14
U-MAE
200
1024
90
16384
61.50
86.04
MAE
641
4096
50
256
51.38
79.60
MAE
641
4096
90
16384
59.12
84.76
Appendix
Table 3: Higher accuracy with longer ViT training schedules. Frozen-probe validation accuracy for dynamic masking at ρ=0.75 , without cropping or flipping during pretraining. Batch sizes are effective batches. All rows use the same initialization; repeated pretraining rows probe the same encoder checkpoint. Schedule names indicate adaptations of the published optimization recipes, with uniformity regularization disabled. Results use one run per configuration and the final probe epoch.
Probe with crop/flip
Probe without crop/flip
Pretraining masks
Top-1
Top-5
Top-1
Top-5
Static
25.34
51.20
26.22
53.28
Dynamic
39.96
68.78
41.98
71.06
Appendix
Table 4: The masking advantage persists without probe augmentation. ViT-Base/ImageNet-100 validation accuracy (%, ↑ ), using the same pretrained encoders and probe hyperparameters in each pair. Both encoders were pretrained without cropping or flipping at ρ=0.75 . Results use one initialization and one probe seed.
Figure 6: Finite-dimensional PCA cosine similarity decreases with dimension. Same bounded-noise setting as Fig. 2 (b): γ=41+49B , B∼Beta(1,2) , and β=1 . (a) Cosine similarity versus sample complexity, with color indicating dimension size d . (b) Dimension dependence at fixed α . Dashed lines fit Aαd−1/2 for d≥1600 , with exponent fixed. Bars show ±1 SD over 30 seeds.
Figure 7: The ViT masking comparison also holds for top-1 accuracy. Frozen linear-probe validation accuracy on ImageNet-100. (a) Static and dynamic masking without cropping or flipping: mean ±1 SD over three seeds. Each horizontal reference uses one seed. (b) Static and dynamic masking with and without cropping and flipping, using the common seed and protocol of Fig. 5 ; no error bars are shown. Training and evaluation settings are given in Sec. H.1 .
Figure 8: Noise-aware spectral baselines for heteroskedastic noise. Cosine similarity versus sample complexity under (a) bounded and (b) unbounded heterogeneous noise. The comparison includes dynamic masking, unmasked reconstruction (ordinary PCA), diagonal-deletion PCA [ 25 ] , HeteroPCA [ 59 ] , estimated-whitening PCA [ 35 ] , and the Bayes-optimal baseline. Lines show theoretical references; markers show numerical estimates. Green curves give the Bayes-optimal state-evolution reference; green diamonds show AMP results ( Secs. F.5 and H.14 ). Bars show ±1 SD; hollow AMP markers identify groups containing iteration-limited runs.
Figure 9: Masking improves recovery across activation functions. Cosine similarity for (a) dynamic masking and (b) no masking in the same setting as Fig. 3 . Colored curves are replica predictions, colored markers in both panels are ERM simulations at finite d=2000 , averaged across 30 seeds ±1 SD. In both panels, the green curve shows the Bayes-optimal prediction and green diamonds show AMP results ( Secs. F.5 and H.14 ).
Figure 10: Mask resampling worsens training error while improving test loss. Centered reconstruction errors versus masking ratio at α=8 under bounded heterogeneous noise. Left: training error. Right: independent test error. Curves are the replica predictions in ( G.3 ) and ( G.4 ); markers and bars show mean ±1 SD over 30 simulations. Lower values are better; zero denotes the zero-decoder baseline. Colors distinguish the fixed mask collections and dynamic masking.
Figure 11: Additional labels resolve the orientation of a frozen feature. Predicted downstream classification accuracy for linear dynamic masking versus sample complexity. The colorbar gives nds on a logarithmic scale. All curves use the same pretraining solution at each α and differ only in the size of the independent labeled dataset used by the scalar logistic probe.
Figure 12: Better reconstruction need not improve CNN representations. (a) Training and validation reconstruction MSE during unmasked CNN pretraining. (b) Linear-probe test accuracy at different pretraining epochs. Shading and bars show ±1 SD over five seeds. The dotted line marks epoch 16 , selected by mean probe validation accuracy for Fig. 1 .
Figure 13: Reconstruction and downstream performance also diverge for ViT. (a) Training and validation reconstruction MSE during unmasked ViT pretraining. (b) Linear-probe top-5 accuracy on the validation set at different pretraining epochs. The dotted line marks six completed pretraining epochs, selected by probe validation accuracy for Fig. 1 . Results use one pretraining seed.
Figure 14: Mask renewal and image augmentation in CNNs. Frozen linear-probe test top-1 accuracy on CIFAR-10 versus masking ratio ρ . Blue/yellow indicate dynamic / static masking; circles/crosses denote pretraining without/with stochastic cropping and flipping. Each marker reports the mean over five seeds, with data splits and autoencoder initializations shared across conditions within each seed. Linear probes use no image augmentation.
Figure 15: CNN training under the original protocol. (a,b) Training and validation reconstruction MSE versus optimizer updates. Thin traces show individual seeds; thick curves and shading show mean ±1 SD while all five runs are observed. Circles and crosses mark early stopped and final checkpoints, respectively. (c) Frozen linear-probe test accuracy of early-stopped encoders, with ±1 SD. The horizontal dashed line indicates dynamic masking.
Figure 16: CNN control with a fixed learning rate and extended training. Panels and uncertainty conventions follow Fig. 15 .
Figure 17: The downstream mask-diversity trend persists across training protocols. Test accuracies from (a) are taken from linear probing the early-stopped checkpoints from Figs. 15 and 16 . (b) corresponds to linear probing the checkpoints after 125 k optimizer updates in the setting of Fig. 16 . Markers and bars show mean ±1 SD over five seeds.
Figure 18: Finite-mask threshold ratios under homogeneous and bounded noise. Rows show (a) homogeneous variance distribution γ=1 and (b) the bounded law ( H.2 ); columns vary the masking ratio ρ . Colors indicate signal strength β . Markers are numerical teacher-overlap instability thresholds divided by their corresponding dynamic threshold ( K→∞ ). The black curve is the weak-signal ratio prediction 1+K−1([2ρ(1−ρ)]−1−1) ; the dashed line marks unity.
Task
Metric
K=1
K=10
Dynamic
MNLI-m
Accuracy
72.85±0.33
76.27±0.42
76.83±0.18
MNLI-mm
Accuracy
74.47±0.29
77.48±0.13
77.65±0.35
QNLI
Accuracy
82.55±0.87
84.40±0.61
85.45±0.38
QQP
Accuracy
85.92±0.10
87.02±0.16
87.22±0.16
SST-2
Accuracy
87.63±0.60
88.08±1.03
89.40±0.46
STS-B
Spearman
79.76±1.18
81.32±1.04
82.46±1.67
Appendix
Table 5: BERT downstream results. Development scores ( ×100 , ↑ ), reported as mean ± SD across three pretraining seeds after averaging two fine-tuning runs per pretraining seed. MCC denotes Matthews correlation; STS-B uses Spearman correlation. MNLI-m and MNLI-mm are the two splits of the same task.