Self-supervised learning (SSL) removes the need for annotations and makes models that are capable across more domains than supervised learning. The autoencoder SSL framework learns by reconstructing its own input after information loss through a bottleneck or noise injection. Masked autoencoders (MAE) are the most successful instantiation of this framework: they encode a random subset of patches, then decode the masked-out patches. In this work, we introduce key modifications to improve MAEs. Our method augments an image in two different ways, then masks and encodes each view separately. It then exchanges the global representations (CLS tokens) between views before decoding the masked patches. By design, our Masked Swingers encourages learning a view-agnostic summary of the image to facilitate efficient transfer. We perform extensive experiments, and find Masked Swingers outperforms MAE by +3-5% on ImageNet-1K kNN and provides large gains on fine-grained tasks, e.g., relative gains of +45% on instance retrieval, +22% on animal re-ID, and +76% on Omniglot character recognition. To boot, Swingers reduces error -64% relative to MAE on three new state-probing datasets, opening the door to world modeling. Welcome to our Swingers party.
Figures & tables
Figure 1: Our Masked Swingers exchanges CLS tokens between different views of the same image prior to reconstruction. We make different views through independent data augmentation and masking. Together, these two design choices encourages the encoder to summarize its input to help reconstruct an arbitrary view so that it can be transferred efficiently and broadly.
Figure 2: Our Half-Swingers extracts better features than MAE evaluated over 10 tasks and 3 pre-training lengths. Our Full-Swingers beats MAE on natural image classification (IN-1K, IN-21K, and mini-VTAB natural) on shorter schedules, and all other tasks on most schedules. In our controlled experiments ( ), we run 9 pre-trainings (varying data augmentation and color masking strengths) per SSL algorithm and pre-training length, choosing the best config per point. Intended for reference only, we also run evaluations on public/official pre-trained checkpoints ( ).
Figure 3: Our Masked Swingers are as fine-tunable as the MAE baselines. In our controlled setting, Half-Swingers ( ) edges Full-Swingers ( ), which edges the MAE/ColorMAE baseline ( ). However, pre-training length affects downstream accuracy much more than these differences. The reported results from He et al. (2022) ( ) and Hinojosa et al. (2024) ( ) outperform these runs, yet they differ with ours in their batch size, learning rate, repeat augmentation, and implementation.
Figure 4: Full-Swingers prefers lower mask rates than Half-Swingers . For example, at a 65% mask rate, our Full-Swingers ties or outperforms Half-Swingers at its best mask rate of 75%. Omniglot is an exception where higher mask rates improves Full-Swingers. Stronger augmentation tends to make Full-Swingers more robust to mask rates and it tends to hurt Half-Swingers.
Figure 5: Full-Swingers and Half-Swingers tend to improve when increasing model size. The gain is monotonic for Omniglot but is more complex for other tasks. For example, Full-Swingers (aug=medium) peaks on animal re-identification at ViT-Small, then decreases for ViT-Base and ViT-Large; increasing augmentation helps for these larger architectures.
Aug.
Color
Pre-training epochs
Strength
Masking
100
200
800
Weak
None
39.1
33.9
33.2
Strong
Red
30.4
37.1
41.2
Table 1: Full-Swingers need stronger augmentation when pre-trained for longer. Results are ImageNet-1K k NN top-1 % accuracy. “Red” masks larger contiguous areas, “None” is uniform sampling as in MAE.
Figure 6: Swingers distributes the variation of the CLS token more evenly across directions. Effective rank of the CLS covariance on IN-1K val vs. (left) 10-shot IN-21K k NN accuracy and (right) Omniglot 20-way one-shot acc., for 81 pre-trained models (3 aug strengths × 3 masking strengths × 3 training lengths × 3 algorithms).
Same Aug.
Swing Strategy
ImageNet-1K
Omniglot
CLS
Pool
FT
CLS
Pool
Yes
None
38.0
37.9
81.4
47.3
40.8
Half
12.1
29.6
81.8
38.5
36.0
Full
11.7
24.1
82.0
43.5
33.8
No
None
35.3
38.1
81.3
41.5
38.3
Half
50.3
37.8
81.5
68.0
37.5
Table 2: Swingers needs independently augmented views to learn “higher-level” frozen features. Encouraging view- specific representations, by sharing augmentations between views, drops IN-1K k NN by −28% and −38%, for Full and Half-Swingers, respectively. Yet it can boost downstream fine-tuning accuracy. “ CLS ” and “Pool” results are computed via k NN. Best and 2 nd best per column.
IN-1K
Omniglot
Setting
CLS
Pool
CLS
Pool
exchange CLS (Full-Swingers)
39.2
35.6
83.5
46.8
exchange CLS on 50% of imgs (Half-Swingers)
50.4
37.8
68.0
37.3
exchange pooled patches
42.8
46.8
60.0
80.5
exchange pooled patches on 50% of imgs
38.4
44.8
51.5
55.0
exchange CLS and pooled patches
36.3
36.8
80.0
64.0
Table 3: Swinging other tokens can also work well, e.g., exchanging both CLS and pooled patches. Exchanging patch-pooled tokens may not need to be half-swung under this setting to perform well (i.e., 47% ImageNet-1k k NN). Best and 2 nd best per column.
Model
Pong
Dino
Golf
Pre-trained by us on ImageNet-1K
MAE/ColorMAE (mean)
0.044
0.058
0.091
MAE/ColorMAE (best)
0.007
0.031
0.033
Half-Swingers (mean)
0.012
0.032
0.054
Half-Swingers (best)
0.005
0.013
0.024
Full-Swingers (mean)
0.009
0.023
0.036
Table 4: Masked Swingers shows promise for world modeling. Masked Swingers improves state probing NMSE (lower ↓ is better) on the three MotionJEPA games ( Karmann et al., 2026 ) . And this improvement is large , e.g., 64% error reduction from MAE/ColorMAE to our Full-Swingers (averaged over all tasks and models). The MotionJEPA result is not directly comparable: its encoder is ViT-T/14 at 1122 px trained from scratch on each game, whereas all other models are ViT-B/16 at 2242 px pre-trained on ImageNet-1K. Best and 2 nd best per column.
Why can masked prediction learn useful representations that unmasked reconstruction misses? We study this question in a high-dimensional model of a masked autoencoder (MAE) trained on data with shared latent structure and heterogeneous noise. We prove that masked linear reconstruction can recover the latent feature at linear sample complexity in regimes where unmasked linear reconstruction, equivalent to PCA, fails. The analysis also quantifies the statistical advantage of mask resampling, an established ingredient of masked pretraining. By introducing a fixed collection of K masks per sample, we characterize its effect on feature recovery and downstream performance, identifying regimes where greater mask diversity lowers sample complexity. Guided by this prediction, we find that random cropping and flipping in standard image-training pipelines can obscure the advantage of mask resampling by renewing the prediction task even when the patch mask is fixed. Removing these transformations reveals a downstream advantage for dynamic over static masking in CNN autoencoders and vision transformers. A complementary BERT pilot finds benefits from greater mask diversity on downstream language tasks. Our results separate the benefit of the masked prediction objective from that of mask diversity, and show how a tractable theory can guide experiments that uncover advantages hidden by standard training practices.
Jorge Medina Moreira, Lorenzo Bardone, Lenka Zdeborová
Statistical Physics of Computation Laboratory, École Polytechnique Fédérale de Lausanne (EPFL), CH-1015 Lausanne, Switzerland
Masked image modeling (MIM) is a highly effective self-supervised learning (SSL) approach to extract useful feature representations from unannotated data. Predominantly used random masking methods make SSL less effective for medical images due to the contextual similarity of neighboring patches, leading to information leakage and SSL simplification. Hierarchical shifted window (Swin) transformer, a highly effective approach for medical images cannot use advanced masking methods as it lacks a global [CLS] token. Hence, we introduced an attention guided masking mechanism for Swin within a co-distillation learning framework to selectively mask semantically co-occurring and discriminative patches, to reduce information leakage and increase the difficulty of SSL pretraining. However, attention guided masking inevitably reduces the diversity of attention heads, which negatively impacts downstream task performance. To address this, we for the first time, integrate a noisy teacher into the co-distillation framework (termed DAGMaN) that performs attentive masking while preserving high attention head diversity. We demonstrate the capability of DAGMaN on multiple tasks including full- and few-shot lung nodule classification, immunotherapy outcome prediction, tumor segmentation, and unsupervised organs clustering.
Deep learning models for medical image classification usually achieve promising results but typically rely on large, annotated datasets or standard transfer learning from ImageNet. Self-Supervised Learning (SSL) has emerged as a powerful alternative, yet common methods like masked autoencoders (MAEs) may inadvertently destroy fine-grained diagnostic features by using random masking. In this paper, we propose a novel SSL pre-training strategy, the Chaotic Denoising Autoencoder (CDAE). Instead of masking, we apply a chaotic transformation to the input image, tasking an autoencoder to reconstruct the original. We hypothesize this forces the encoder to learn robust, domain-specific features by "inverting the chaos". Furthermore, we propose an attentive fusion mechanism that combines features from our CDAE-trained encoder with a standard encoder, leveraging the strengths of both general and domain-specific representations. Our method is evaluated on two public medical datasets: ISIC 2018 (skin lesions) and APTOS 2019 (diabetic retinopathy). The proposed model achieves high performance, with an accuracy of 0.9221 and an F1-macro of 0.8530 on ISIC 2018, and an accuracy of 0.8644 and F1-macro of 0.7433 on APTOS 2019, demonstrating the efficacy of our approach.
Joao Batista Florindo, Amanda Pontes de Oliveira Ornelas
Institute of Mathematics, Statistics and Scientific Computing - University of Campinas, Rua S´ergio Buarque de Holanda, 651, Campinas, Brasil