cs.LGMar 20, 2025

Structured-Noise Masked Modeling for Video, Audio and Beyond

Authors: Aritra Bhowmik, Carlos Hinojosa, Fida Mohammad Thoker, Bernard Ghanem, Cees G. M. Snoek

Organizations: University of Amsterdam · King Abdullah University of Science and Technology

Abstract

Masked modeling has emerged as a robust self-supervised learning framework. However, most methods rely on random masking, which disregards the structural properties of different data modalities. To align with the spatiotemporal and spectral characteristics of video and audio data, we introduce a structured noise-based masking approach. By filtering white noise into different color noise distributions, we generate structured masks that capture modality-specific patterns without requiring handcrafted heuristics or access to the data. Our approach enhances masked video and audio modeling frameworks without any additional computational cost. Experiments show that structured noise masking consistently outperforms random masking, underscoring the value of modality-aware masking strategies for representation learning.

Figures & tables

Appendix figures & tables20 assets

Supplementary material from the paper’s appendix.

Appendix

Explore similar work

CardsList
  1. Image Classifiers are Efficient Self-Supervised Video Representation Learners

    Sep 30, 2026Owais Iqbal, Sudipta Sarkar, Shyam Marjit +3Vision TransformerImage Classification

  2. AudioMosaic: Contrastive Masked Audio Representation Learning

    May 14, 2026Hanxun Huang, Qizhou Wang, Xingjun Ma +3Neural AudioOpencode

  3. The hidden advantage of mask resampling: a theory of masked autoencoders

    Oct 1, 2026Jorge Medina Moreira, Lorenzo Bardone, Lenka Zdeborová3D Masked AutoencodersAutoencoder Architectures