cs.LGSep 29, 2026

Predictive Self-Supervised Learning Provably Identifies Stochastic Signals under Nuisance

Authors: Fabian A. Mikulasch, Friedemann Zenke

Organizations: Friedrich Miescher Institute for Biomedical Research, Basel, Switzerland · Faculty of Science, University of Basel, Switzerland

Abstract

Self-supervised learning (SSL) by predicting in latent space, without generating the input data itself, learns highly abstract, useful representations. Intuitively, this success is often attributed to its ability to discard nuisance information that is irrelevant to prediction. However, this poses a conundrum: both stochastic variation in a prediction-relevant latent signal and true nuisance make observations partly unpredictable; how could they be distinguished? Surprisingly, we prove that common SSL methods can achieve exactly this, by implicitly instantiating a latent-variable model with stochastic dynamics and observation-private nuisance. We trace their ability to recover the stochastic signal to two complementary principles: Predictive mutual information maximization ensures that representations retain the information needed for prediction, while latent distribution matching constrains how this information is encoded, thereby making the retained signal identifiable. We confirm this identifiability result in simulations for Gaussian predictors, which recover the true signal up to an affine transformation even in dynamic, nuisance-laden environments.

Figures & tables

Appendix figures & tables2 assets

Supplementary material from the paper’s appendix.

Appendix

Explore similar work

CardsList
  1. Understanding Self-Supervised Learning via Latent Distribution Matching

    May 5, 2026Fabian A Mikulasch, Friedemann ZenkeFuture Latent RepresentationsBayesian Filtering

  2. Martingale-Consistent Self-Supervised Learning

    May 12, 2026Moritz Gögl, Hanwen Xing, Christopher YauSelf-Supervised LearningConformal Test Martingales

  3. Self-Supervised Learning from Structural Invariance

    Feb 2, 2026Yipeng Zhang, Hafez Ghaemi, Jungyoon Lee +3Self-Supervised LearningLatent Representation Alignment