Self-supervised learning (SSL) removes the need for annotations and makes models that are capable across more domains than supervised learning. The autoencoder SSL framework learns by reconstructing its own input after information loss through a bottleneck or noise injection. Masked autoencoders (MAE) are the most successful instantiation of this framework: they encode a random subset of patches, then decode the masked-out patches. In this work, we introduce key modifications to improve MAEs. Our method augments an image in two different ways, then masks and encodes each view separately. It then exchanges the global representations (CLS tokens) between views before decoding the masked patches. By design, our Masked Swingers encourages learning a view-agnostic summary of the image to facilitate efficient transfer. We perform extensive experiments, and find Masked Swingers outperforms MAE by +3-5% on ImageNet-1K kNN and provides large gains on fine-grained tasks, e.g., relative gains of +45% on instance retrieval, +22% on animal re-ID, and +76% on Omniglot character recognition. To boot, Swingers reduces error -64% relative to MAE on three new state-probing datasets, opening the door to world modeling. Welcome to our Swingers party.
Figures & tables
Figure 1: Our Masked Swingers exchanges CLS tokens between different views of the same image prior to reconstruction. We make different views through independent data augmentation and masking. Together, these two design choices encourages the encoder to summarize its input to help reconstruct an arbitrary view so that it can be transferred efficiently and broadly.
Figure 2: Our Half-Swingers extracts better features than MAE evaluated over 10 tasks and 3 pre-training lengths. Our Full-Swingers beats MAE on natural image classification (IN-1K, IN-21K, and mini-VTAB natural) on shorter schedules, and all other tasks on most schedules. In our controlled experiments ( ), we run 9 pre-trainings (varying data augmentation and color masking strengths) per SSL algorithm and pre-training length, choosing the best config per point. Intended for reference only, we also run evaluations on public/official pre-trained checkpoints ( ).
Figure 3: Our Masked Swingers are as fine-tunable as the MAE baselines. In our controlled setting, Half-Swingers ( ) edges Full-Swingers ( ), which edges the MAE/ColorMAE baseline ( ). However, pre-training length affects downstream accuracy much more than these differences. The reported results from He et al. (2022) ( ) and Hinojosa et al. (2024) ( ) outperform these runs, yet they differ with ours in their batch size, learning rate, repeat augmentation, and implementation.
Figure 4: Full-Swingers prefers lower mask rates than Half-Swingers . For example, at a 65% mask rate, our Full-Swingers ties or outperforms Half-Swingers at its best mask rate of 75%. Omniglot is an exception where higher mask rates improves Full-Swingers. Stronger augmentation tends to make Full-Swingers more robust to mask rates and it tends to hurt Half-Swingers.
Figure 5: Full-Swingers and Half-Swingers tend to improve when increasing model size. The gain is monotonic for Omniglot but is more complex for other tasks. For example, Full-Swingers (aug=medium) peaks on animal re-identification at ViT-Small, then decreases for ViT-Base and ViT-Large; increasing augmentation helps for these larger architectures.
Aug.
Color
Pre-training epochs
Strength
Masking
100
200
800
Weak
None
39.1
33.9
33.2
Strong
Red
30.4
37.1
41.2
Table 1: Full-Swingers need stronger augmentation when pre-trained for longer. Results are ImageNet-1K k NN top-1 % accuracy. “Red” masks larger contiguous areas, “None” is uniform sampling as in MAE.
Figure 6: Swingers distributes the variation of the CLS token more evenly across directions. Effective rank of the CLS covariance on IN-1K val vs. (left) 10-shot IN-21K k NN accuracy and (right) Omniglot 20-way one-shot acc., for 81 pre-trained models (3 aug strengths × 3 masking strengths × 3 training lengths × 3 algorithms).
Same Aug.
Swing Strategy
ImageNet-1K
Omniglot
CLS
Pool
FT
CLS
Pool
Yes
None
38.0
37.9
81.4
47.3
40.8
Half
12.1
29.6
81.8
38.5
36.0
Full
11.7
24.1
82.0
43.5
33.8
No
None
35.3
38.1
81.3
41.5
38.3
Half
50.3
37.8
81.5
68.0
37.5
Table 2: Swingers needs independently augmented views to learn “higher-level” frozen features. Encouraging view- specific representations, by sharing augmentations between views, drops IN-1K k NN by −28% and −38%, for Full and Half-Swingers, respectively. Yet it can boost downstream fine-tuning accuracy. “ CLS ” and “Pool” results are computed via k NN. Best and 2 nd best per column.
IN-1K
Omniglot
Setting
CLS
Pool
CLS
Pool
exchange CLS (Full-Swingers)
39.2
35.6
83.5
46.8
exchange CLS on 50% of imgs (Half-Swingers)
50.4
37.8
68.0
37.3
exchange pooled patches
42.8
46.8
60.0
80.5
exchange pooled patches on 50% of imgs
38.4
44.8
51.5
55.0
exchange CLS and pooled patches
36.3
36.8
80.0
64.0
Table 3: Swinging other tokens can also work well, e.g., exchanging both CLS and pooled patches. Exchanging patch-pooled tokens may not need to be half-swung under this setting to perform well (i.e., 47% ImageNet-1k k NN). Best and 2 nd best per column.
Model
Pong
Dino
Golf
Pre-trained by us on ImageNet-1K
MAE/ColorMAE (mean)
0.044
0.058
0.091
MAE/ColorMAE (best)
0.007
0.031
0.033
Half-Swingers (mean)
0.012
0.032
0.054
Half-Swingers (best)
0.005
0.013
0.024
Full-Swingers (mean)
0.009
0.023
0.036
Table 4: Masked Swingers shows promise for world modeling. Masked Swingers improves state probing NMSE (lower ↓ is better) on the three MotionJEPA games ( Karmann et al., 2026 ) . And this improvement is large , e.g., 64% error reduction from MAE/ColorMAE to our Full-Swingers (averaged over all tasks and models). The MotionJEPA result is not directly comparable: its encoder is ViT-T/14 at 1122 px trained from scratch on each game, whereas all other models are ViT-B/16 at 2242 px pre-trained on ImageNet-1K. Best and 2 nd best per column.