Masked modeling has emerged as a robust self-supervised learning framework. However, most methods rely on random masking, which disregards the structural properties of different data modalities. To align with the spatiotemporal and spectral characteristics of video and audio data, we introduce a structured noise-based masking approach. By filtering white noise into different color noise distributions, we generate structured masks that capture modality-specific patterns without requiring handcrafted heuristics or access to the data. Our approach enhances masked video and audio modeling frameworks without any additional computational cost. Experiments show that structured noise masking consistently outperforms random masking, underscoring the value of modality-aware masking strategies for representation learning.
Figures & tables
Figure 1
Figure 3 : a Illustration of the metric used to determine the concentration of visible patches in a window UPi of the mask Mx,yi . b Example of the initial mask ( Mi ), with clusters of visible patches, and c final mask ( M^bi ) obtained with our regularized blue noise masking algorithm, with uniformly distributed visible patches. Note the improved uniformity in the final mask, which ensures better coverage and reduces undesirable clustering effects.
Masking Type
SSv2 Pretraining
K400 Pretraining
Method
Data-independant
Data-adaptive
SSv2 Top-1
SSv2 Top-1
K400 Top-1
VideoMAE
Random
-
69.6
68.5
80.0
VideoMAE + Our masking
Green3D
-
70.8 (+1.2%)
69.7 (+1.2%)
80.5 (+0.5%)
CMAE-V
Random
-
69.7
-
80.2
OmniMAE
Random
-
69.5
69.0
80.8
MME
Random
-
70.0
70.5
81.5
Table 1 : Comparison of masked video modeling methods on Something-Something V2 and Kinetics-400 for standard action recognition. All results use a ViT-B backbone pretrained on K400 or SSv2 for 800 epochs. Our Green3D masking consistently improves both in-domain and cross-domain performance for VideoMAE as well as advanced architectures such as SIGMA.
Clustering
Overclustering
Method
YTVOS
DAVIS
YTVOS
DAVIS
VideoMAE
34.1
29.5
61.3
56.2
VideoMAE + Our masking
35.6 (+1.5%)
38.2 (+8.7%)
62.5 (+1.5%)
58.2 (+2.0%)
MGM
36.6
36.5
61.2
56.6
MGMAE
34.5
31.0
60.1
57.5
SIGMA
41.1
33.1
67.1
59.0
Table 2 : Comparison of masked video methods for unsupervised video object segmentation. Following the evaluation protocol from [ 43 ] , we report mIoU for clustering and overclustering. We evaluate the ViT-B backbone pretrained on K400 and use the officially released checkpoints for all prior works. Equipping VideoMAE with our masking significantly improves its performance, surpassing even motion-guided masking methods. Our masking also improves SIGMA when used as a plugin.
Domain
Sample
Action
Task
Method
shift
efficiency
granularity
shift
Mean
SSv2
Gym99
UCF( 103 )
GYM( 103 )
FX-S1
UB-S1
UCF-RC ↓
Charades
VideoMAE
68.6
86.6
74.6
25.9
36.6
74.3
0.172
17.2
58.3
VideoMAE + Our masking
69.7
88.1
75.0
29.9
38.6
74.8
0.170
18.2
59.6
MVD
70.0
82.5
67.1
17.5
31.3
50.5
0.184
16.1
52.1
MGMAE
68.9
87.2
77.2
24.1
33.7
79.5
0.181
17.9
58.8
Table 3 : Comparison on the SEVERE benchmark [ 51 ] evaluating domain shift, sample efficiency, action granularity, and task shift of learned video representations. All methods are pretrained on K400 with a ViT-B backbone. Our method improves the downstream generalization of both VideoMAE and SIGMA architectures.
Table 6Table 7
Method
SEVERE
Clustering
Overclustering
Gym99
FX-S1
UB-S1
Charades
YTVOS
DAVIS
YTVOS
DAVIS
Random
72.0
35.7
71.3
14.2
29.7
25.3
56.2
43.9
Red3D
72.5
37.5
73.0
15.4
31.2
25.9
54.7
48.6
Blue3D
73.4
34.7
69.3
15.8
31.3
25.5
55.6
44.1
Green3D
75.4
38.5
72.6
16.0
32.9
27.3
60.0
50.9
Table 8 : SEVERE fine-grained and segmentation results for 3D masking. Compared with the random baseline and other 3D masks, Green3D achieves the strongest overall performance across selected SEVERE tasks and unsupervised video object segmentation.
Figure 4 : Qualitative results on DAVIS. Green3D masking produces sharper and temporally consistent segmentations compared to VideoMAE.
Appendix figures & tables20 assets
Supplementary material from the paper’s appendix.
Appendix
Variant
mini-Kinetics
mini-SSv2
Variant-1
52.3
54.3
Variant-2
52.1
53.3
Variant-3
52.2
54.4
Variant-4
51.8
54.3
Variant-5
52.7
54.5
Appendix
Table 0.A.1 : Ablation on σ1 and σ2 in Green 3D noise. Selecting σ values from a controlled range (Variant-5) achieves the best performance, balancing spatial coherence and temporal smoothness.
Window sizes (Δ)
AS-20k
ESC-50
(1,3)
36.3
93.8
(3,5)
36.8
94.6
(7,9)
36.1
93.7
Appendix
Table 0.A.2 : Impact of window size Δ on R-BN masking within AudioMAE. The default multi–scale configuration (3,5) performs best, while all settings outperform vanilla blue noise.
Masking ratio
L2-loss
mini-Kinetics
mini-SSv2
80%
0.48
51.6
53.8
85%
0.53
52.4
54.4
90%
0.60
52.7
54.5
Appendix
Table 0.A.3 : Impact of masking ratio for VideoMAE (Green3D noise). The standard ratio of 90% yields the best performance.
Masking ratio
L2-loss
AS-20k
ESC-50
75%
0.47
36.4
93.9
80%
0.49
36.8
94.6
85%
0.53
36.3
93.4
Appendix
Table 0.A.4 : Impact of masking ratio for AudioMAE (Regularized Blue noise). The standard ratio of 80% performs optimally.
Masking type
mini-Kinetics
mini-SSv2
Grid
51.0
52.3
Block
51.1
52.5
Tube
51.6
52.8
Green3D
52.7
54.5
Appendix
Table 0.A.5 : Comparison of Green3D with common non-learned structured masking strategies. Grid, block, and tube masks underperform Green3D, highlighting the importance of modality-aware mask design for video masked modeling.
Masking
Params
Flops
Mem
Epoch Time
Model
Strategy
(M)
(G)
(GB)
(mm:ss)
VideoMAE
Random
94.21
21.28
22.62
15:18
VideoMAE+ our masking
Green3D
94.21
21.02
22.68
15:31
Appendix
Table 0.A.6 : Computational efficiency of VideoMAE with random masking and with Green3D masking on K400 (ViT-B). Both variants show identical parameters, FLOPs, and memory usage, with only a marginal difference in training time per epoch.
Method
200
400
600
700
800
VideoMAE
66.71
67.76
68.15
68.51
68.53
VideoMAE + our Green3D
67.42
68.45
69.48
69.74
69.76
Appendix
Table 0.A.7 : Convergence comparison on SSv2 with K400 pretraining. Green3D achieves consistent accuracy gains over VideoMAE at all epochs, with improvements increasing from +0.7% at 200 epochs to +1.2% at 600 epochs, while maintaining a similar convergence speed.
Green3D mask (seed)
mini-SSv2
mini-Kinetics
Seed A
54.3
52.6
Seed B
54.5
52.8
Seed C
54.5
52.7
Mean
54.43
52.7
Appendix
Table 0.A.8 : Performance variance across different Green3D mask realizations generated with independent random seeds. Variability remains below 0.2% in absolute terms, indicating that improvements are not driven by stochastic noise patterns.
config
SSv2
K400
optimizer
AdamW
base learning rate
1.5e-4
weight decay
0.05
optimizer momentum
β1,β2=0.9,0.95
batch size
256
learning rate schedule
cosine decay
Appendix
Table 0.A.9: VideoMAE and SIGMA pretraining setup.
config
SSv2
K400
SEVERE
optimizer
AdamW
base learning rate
1.0e-3
weight decay
0.05
optimizer momentum
β1,β2=0.9,0.999
layer-wise lr decay [ 7 ]
0.75
batch size
32
16
16
Appendix
Table 0.A.10: VideoMAE and SIGMA fine-tuning setup.
Evaluation Setup
Experiment
Dataset
Task
#Classes
#Finetuning
#Testing
Eval Metric
Gym99
FineGym
Action Class.
99
20,484
8,521
Top-1 Acc.
Sample Efficiency
UCF ( 103 )
UCF 101
Action Class.
101
1,000
3,783
Top-1 Acc.
Gym ( 103 )
FineGym
Action Class.
99
1,000
8,521
Top-1 Acc.
Action Granularity
FX-S1
FineGym
Action Class.
11
1,882
777
Mean-per-class
UB-S1
FineGym
Action Class.
15
3,511
1,471
Mean-per-class
Task Shift
UCF-RC
UCFRep
Repetition Counting
-
421
105
Mean Error
Appendix
Table 0.A.11 : Details of evaluation subsets in SEVERE benchmark [ 51 ]
Config
Value
Optimizer
AdamW
Base learning rate
2e-4
Weight decay
0.05
Optimizer momentum
β1,β2=0.9,0.95
Batch size
512
Learning rate schedule
Cosine decay
Appendix
Table 0.A.12: AudioMAE pretraining setup.
Config
Value
Optimizer
AdamW
Base learning rate
1e-3
Weight decay
0.05
Optimizer momentum
β1,β2=0.9,0.999
Batch size
256
Learning rate schedule
Cosine decay
Appendix
Table 0.A.13: AudioMAE fine-tuning setup.
Configuration
VGGSound
Optimizer
AdamW
Base learning rate
1e-4
Weight decay
5e-7
Optimizer momentum
β1,β2=0.95,0.999
Batch size
120
Learning rate schedule
Cosine decay
Appendix
Table 0.A.14: CAV-MAE pretraining setup.
Configuration
VGGSound
Optimizer
AdamW
Base learning rate
1e-4
Weight decay
0.05
Batch size
48
Learning rate schedule
Cosine decay
Warmup epochs
2
Appendix
Table 0.A.15: CAV-MAE fine-tuning.
Figure 0.A.1 : Comparison of different masking strategies in VideoMAE pretraining on SSv2 videos (masking ratio 0.75). Standard tube masking struggles to align with video structures, while 2D noise-based masking introduces some spatial coherence but lacks temporal consistency. Our proposed Green3D masking effectively captures spatiotemporal structures, preserving motion continuity across frames.
Figure 0.A.2 : Comparison of different masking strategies in VideoMAE pretraining on SSv2 videos (masking ratio 0.9). Standard tube masking struggles to align with video structures, while 2D noise-based masking introduces some spatial coherence but lacks temporal consistency. Our proposed Green3D masking effectively captures spatiotemporal structures, preserving motion continuity across frames.
Figure 0.A.3 : Comparison of different masking strategies in AudioMAE pretraining on spectrograms (masking ratio 0.8). Random masking leads to scattered reconstructions, whereas red and green noise masking introduce biases that distort the frequency structure. Our proposed Regularized Blue noise masking ensures a more balanced reconstruction by aligning with the spectral distribution of audio signals.
Figure 0.A.4 : Comparison of different masking strategies in AudioMAE pretraining on spectrograms (masking ratio 0.8). Random masking leads to scattered reconstructions, whereas red and green noise masking introduce biases that distort the frequency structure. Our proposed Regularized Blue noise masking ensures a more balanced reconstruction by aligning with the spectral distribution of audio signals.
Figure 0.A.5 : Frequency-domain analysis for (a) audio spectrograms and (b) video frames, comparing the modality spectrum with four masking strategies (Red, Green, Blue, Random).
We introduce VideoMSN, a Masked Siamese Network framework for efficient self-supervised spatio-temporal representation learning in videos. Instead of relying on heavy 3D architectures or reconstruction-based autoencoders for learning with unlabeled data, we repurpose standard image Vision Transformers by representing videos as super images which are grids composed of frames sampled from videos. From each super image, we construct two views: one with spatial patch masking and the other with temporal frame masking, ensuring no information leakage across frames. A shared Vision Transformer (ViT) encoder aligns their embeddings using a masked Siamese loss, capturing both motion and appearance cues without reconstruction. Our decoder-free formulation leverages an image foundation model towards efficient video representation learning. Starting from pretrained DINO-v3 and DeiT-v3 image encoders, VideoMSN achieves state-of-the-art performance on Kinetics-400, UCF101, and HMDB51 while requiring up to 32× fewer and 160× fewer video pretraining epochs compared to prior video self-supervised learning methods. Our proposed approach also shows strong performance in low-shot classification, confirming the transferability of the learned representations in a label-scarce scenario. Project Page: https://cvir.github.io/projects/videomsn.
Owais Iqbal, Sudipta Sarkar, Shyam Marjit +3
Indian Institute of Technology Kharagpur, India · Indian Institute of Science Bangalore, India · École de technologie supérieure Montreal, Canada
Audio self-supervised learning (SSL) aims to learn general-purpose representations from large-scale unlabeled audio data. While recent advances have been driven mainly by generative reconstruction objectives, contrastive approaches remain less explored, partly due to the difficulty of designing effective audio augmentations and the large batch sizes required for contrastive pre-training. We introduce \textbf{AudioMosaic}, a contrastive learning-based audio encoder for general audio understanding. During pre-training, AudioMosaic constructs positive pairs by applying structured time-frequency masking to spectrogram patches, which reduces memory usage and enables efficient large-batch training. Compared with generative approaches, the AudioMosaic encoder learns more discriminative utterance-level representations that demonstrate strong transferability across datasets, domains, and acoustic conditions. Extensive experiments show that AudioMosaic achieves state-of-the-art performance on several standard audio benchmarks under both linear probing and fine-tuning. We further show that integrating the pretrained AudioMosaic encoder into audio-language models improves performance on audio-language tasks. The code is publicly available in our \href{https://github.com/HanxunH/AudioMosaic}{GitHub repository}.
Hanxun Huang, Qizhou Wang, Xingjun Ma +3
School of Computing and Information Systems, The University of Melbourne, Australia · Institute of Trustworthy Embodied AI, Fudan University, China · Baskin School of Engineering, University of California, Santa Cruz, USA
Why can masked prediction learn useful representations that unmasked reconstruction misses? We study this question in a high-dimensional model of a masked autoencoder (MAE) trained on data with shared latent structure and heterogeneous noise. We prove that masked linear reconstruction can recover the latent feature at linear sample complexity in regimes where unmasked linear reconstruction, equivalent to PCA, fails. The analysis also quantifies the statistical advantage of mask resampling, an established ingredient of masked pretraining. By introducing a fixed collection of K masks per sample, we characterize its effect on feature recovery and downstream performance, identifying regimes where greater mask diversity lowers sample complexity. Guided by this prediction, we find that random cropping and flipping in standard image-training pipelines can obscure the advantage of mask resampling by renewing the prediction task even when the patch mask is fixed. Removing these transformations reveals a downstream advantage for dynamic over static masking in CNN autoencoders and vision transformers. A complementary BERT pilot finds benefits from greater mask diversity on downstream language tasks. Our results separate the benefit of the masked prediction objective from that of mask diversity, and show how a tractable theory can guide experiments that uncover advantages hidden by standard training practices.
Jorge Medina Moreira, Lorenzo Bardone, Lenka Zdeborová
Statistical Physics of Computation Laboratory, École Polytechnique Fédérale de Lausanne (EPFL), CH-1015 Lausanne, Switzerland