cs.CVSep 25, 2026

Conditional Predictive Sufficient Statistics for Visual Representation Learning

Authors: Yuzhou Hong

Organizations: Zhejiang Sci-Tech University

Abstract

A useful visual representation is a statistic of the observed past that retains the latent factors shared with the future and discards patch-private noise. We formalize this requirement as a conditional predictive sufficient statistic (CPSS). Under a shared-factor model of image patches, the mutual information between the past and the next patch equals the information the past carries about the shared factor, up to a remainder that the next patch itself fails to reveal. Predicting the next patch embedding with a cosine loss is maximum likelihood for a von Mises-Fisher model of that embedding's direction, and is therefore a tractable surrogate for the predictive information. The same population loss is also minimized by a constant embedding, so stop-gradient does not by itself select the sufficient statistic; it only blocks the symmetric gradient that implements the constant solution in one step. The regression target is a shallow embedding, which forces the network output back into that shallow range and leaves the sufficient statistic in intermediate blocks. Small causal Transformers on MNIST and CIFAR-10 are used as diagnostics, not as a leaderboard. On MNIST the future shift and the stop-gradient move probe accuracy by tens of points, and the CPSS readout peaks before the output. On CIFAR-10, with the same short budget and no augmentation, every objective lands near a linear classifier on pixels. What still matches the derivation is the geometry: the CPSS output is a worse readout than its best intermediate block, next-pixel regression does not pay that penalty, and removing the stop-gradient collapses the effective rank of the embedding even when the pretext loss looks perfect.

Figures & tables

Appendix figures & tables3 assets

Supplementary material from the paper’s appendix.

Appendix

Explore similar work

CardsList
  1. Three Necessary Principles for Self-Supervised Visual Representation Learning

    Aug 8, 2026Nikos Giakoumoglou, Paschalis Giakoumoglou, Tania StathakiSelf-Supervised RepresentationsSelf-Supervised Learning

  2. Generative Residual Factorization

    Sep 28, 2026Letian Gong, Yuzhou HongExploratory Factor Analysis

  3. PEIRA: Learning Predictive Encoders through Inter-View Regressor Alignment

    May 17, 2026Michael Arbel, Basile Terver, Jean PonceContrastive LearningEquilibrium