Structured-Noise Masked Modeling for Video, Audio and Beyond
Organizations: University of Amsterdam · King Abdullah University of Science and Technology
Abstract
Masked modeling has emerged as a robust self-supervised learning framework. However, most methods rely on random masking, which disregards the structural properties of different data modalities. To align with the spatiotemporal and spectral characteristics of video and audio data, we introduce a structured noise-based masking approach. By filtering white noise into different color noise distributions, we generate structured masks that capture modality-specific patterns without requiring handcrafted heuristics or access to the data. Our approach enhances masked video and audio modeling frameworks without any additional computational cost. Experiments show that structured noise masking consistently outperforms random masking, underscoring the value of modality-aware masking strategies for representation learning.
Figures & tables
| Masking Type | SSv2 Pretraining | K400 Pretraining | |||
|---|---|---|---|---|---|
| Method | Data-independant | Data-adaptive | SSv2 Top-1 | SSv2 Top-1 | K400 Top-1 |
| VideoMAE | Random | - | 69.6 | 68.5 | 80.0 |
| VideoMAE + Our masking | Green3D | - | 70.8 (+1.2%) | 69.7 (+1.2%) | 80.5 (+0.5%) |
| CMAE-V | Random | - | 69.7 | - | 80.2 |
| OmniMAE | Random | - | 69.5 | 69.0 | 80.8 |
| MME | Random | - | 70.0 | 70.5 | 81.5 |
| Clustering | Overclustering | |||
|---|---|---|---|---|
| Method | YTVOS | DAVIS | YTVOS | DAVIS |
| VideoMAE | 34.1 | 29.5 | 61.3 | 56.2 |
| VideoMAE + Our masking | 35.6 (+1.5%) | 38.2 (+8.7%) | 62.5 (+1.5%) | 58.2 (+2.0%) |
| MGM | 36.6 | 36.5 | 61.2 | 56.6 |
| MGMAE | 34.5 | 31.0 | 60.1 | 57.5 |
| SIGMA | 41.1 | 33.1 | 67.1 | 59.0 |
| Domain | Sample | Action | Task | ||||||
|---|---|---|---|---|---|---|---|---|---|
| Method | shift | efficiency | granularity | shift | Mean | ||||
| SSv2 | Gym99 | UCF( ) | GYM( ) | FX-S1 | UB-S1 | UCF-RC | Charades | ||
| VideoMAE | 68.6 | 86.6 | 74.6 | 25.9 | 36.6 | 74.3 | 0.172 | 17.2 | 58.3 |
| VideoMAE + Our masking | 69.7 | 88.1 | 75.0 | 29.9 | 38.6 | 74.8 | 0.170 | 18.2 | 59.6 |
| MVD | 70.0 | 82.5 | 67.1 | 17.5 | 31.3 | 50.5 | 0.184 | 16.1 | 52.1 |
| MGMAE | 68.9 | 87.2 | 77.2 | 24.1 | 33.7 | 79.5 | 0.181 | 17.9 | 58.8 |
| Method | SEVERE | Clustering | Overclustering | |||||
|---|---|---|---|---|---|---|---|---|
| Gym99 | FX-S1 | UB-S1 | Charades | YTVOS | DAVIS | YTVOS | DAVIS | |
| Random | 72.0 | 35.7 | 71.3 | 14.2 | 29.7 | 25.3 | 56.2 | 43.9 |
| Red3D | 72.5 | 37.5 | 73.0 | 15.4 | 31.2 | 25.9 | 54.7 | 48.6 |
| Blue3D | 73.4 | 34.7 | 69.3 | 15.8 | 31.3 | 25.5 | 55.6 | 44.1 |
| Green3D | 75.4 | 38.5 | 72.6 | 16.0 | 32.9 | 27.3 | 60.0 | 50.9 |
Appendix figures & tables20 assets
Supplementary material from the paper’s appendix.
Appendix
| Variant | mini-Kinetics | mini-SSv2 |
|---|---|---|
| Variant-1 | 52.3 | 54.3 |
| Variant-2 | 52.1 | 53.3 |
| Variant-3 | 52.2 | 54.4 |
| Variant-4 | 51.8 | 54.3 |
| Variant-5 | 52.7 | 54.5 |
| Window sizes | AS-20k | ESC-50 |
|---|---|---|
| (1,3) | 36.3 | 93.8 |
| (3,5) | 36.8 | 94.6 |
| (7,9) | 36.1 | 93.7 |
| Masking ratio | L2-loss | mini-Kinetics | mini-SSv2 |
|---|---|---|---|
| 80% | 0.48 | 51.6 | 53.8 |
| 85% | 0.53 | 52.4 | 54.4 |
| 90% | 0.60 | 52.7 | 54.5 |
| Masking ratio | L2-loss | AS-20k | ESC-50 |
|---|---|---|---|
| 75% | 0.47 | 36.4 | 93.9 |
| 80% | 0.49 | 36.8 | 94.6 |
| 85% | 0.53 | 36.3 | 93.4 |
| Masking type | mini-Kinetics | mini-SSv2 |
|---|---|---|
| Grid | 51.0 | 52.3 |
| Block | 51.1 | 52.5 |
| Tube | 51.6 | 52.8 |
| Green3D | 52.7 | 54.5 |
| Masking | Params | Flops | Mem | Epoch Time | |
|---|---|---|---|---|---|
| Model | Strategy | (M) | (G) | (GB) | (mm:ss) |
| VideoMAE | Random | 94.21 | 21.28 | 22.62 | 15:18 |
| VideoMAE+ our masking | Green3D | 94.21 | 21.02 | 22.68 | 15:31 |
| Method | 200 | 400 | 600 | 700 | 800 |
|---|---|---|---|---|---|
| VideoMAE | 66.71 | 67.76 | 68.15 | 68.51 | 68.53 |
| VideoMAE + our Green3D | 67.42 | 68.45 | 69.48 | 69.74 | 69.76 |
| Green3D mask (seed) | mini-SSv2 | mini-Kinetics |
|---|---|---|
| Seed A | 54.3 | 52.6 |
| Seed B | 54.5 | 52.8 |
| Seed C | 54.5 | 52.7 |
| Mean | 54.43 | 52.7 |
| config | SSv2 | K400 |
|---|---|---|
| optimizer | AdamW | |
| base learning rate | 1.5e-4 | |
| weight decay | 0.05 | |
| optimizer momentum | ||
| batch size | 256 | |
| learning rate schedule | cosine decay | |
| config | SSv2 | K400 | SEVERE |
|---|---|---|---|
| optimizer | AdamW | ||
| base learning rate | 1.0e-3 | ||
| weight decay | 0.05 | ||
| optimizer momentum | |||
| layer-wise lr decay [ 7 ] | 0.75 | ||
| batch size | 32 | 16 | 16 |
| Evaluation Setup | Experiment | Dataset | Task | #Classes | #Finetuning | #Testing | Eval Metric |
|---|---|---|---|---|---|---|---|
| Gym99 | FineGym | Action Class. | 99 | 20,484 | 8,521 | Top-1 Acc. | |
| Sample Efficiency | UCF ( ) | UCF 101 | Action Class. | 101 | 1,000 | 3,783 | Top-1 Acc. |
| Gym ( ) | FineGym | Action Class. | 99 | 1,000 | 8,521 | Top-1 Acc. | |
| Action Granularity | FX-S1 | FineGym | Action Class. | 11 | 1,882 | 777 | Mean-per-class |
| UB-S1 | FineGym | Action Class. | 15 | 3,511 | 1,471 | Mean-per-class | |
| Task Shift | UCF-RC | UCFRep | Repetition Counting | - | 421 | 105 | Mean Error |
| Config | Value |
|---|---|
| Optimizer | AdamW |
| Base learning rate | 2e-4 |
| Weight decay | 0.05 |
| Optimizer momentum | |
| Batch size | 512 |
| Learning rate schedule | Cosine decay |
| Config | Value |
|---|---|
| Optimizer | AdamW |
| Base learning rate | 1e-3 |
| Weight decay | 0.05 |
| Optimizer momentum | |
| Batch size | 256 |
| Learning rate schedule | Cosine decay |
| Configuration | VGGSound |
|---|---|
| Optimizer | AdamW |
| Base learning rate | 1e-4 |
| Weight decay | 5e-7 |
| Optimizer momentum | |
| Batch size | 120 |
| Learning rate schedule | Cosine decay |
| Configuration | VGGSound |
|---|---|
| Optimizer | AdamW |
| Base learning rate | 1e-4 |
| Weight decay | 0.05 |
| Batch size | 48 |
| Learning rate schedule | Cosine decay |
| Warmup epochs | 2 |