cs.LGSep 29, 2026

Which Tasks Survive Self-Supervised Learning?

Authors: Achleshwar Luthra, Lucas Bryant, Tracy Zhu, Tomer Galanti

Organizations: Department of Computer Science and Engineering Texas A&M University

Abstract

Same-instance self-supervised learning (SSL) learns representations by enforcing consistency across two views of the same underlying instance. This principle alone, however, does not determine which downstream tasks remain recoverable from the learned representation. We study this question through \emph{semantic recoverability}, defined as the amount of a task's posterior score captured by the represented function space. We show that, for centered and whitened representations, recoverability exactly determines directional class-distance-normalized variance (CDNV), controls few-shot nearest-centroid classification, and governs the strength of task-relevant semantic directions. The population linear probe and centroid axis coincide, and multiple well-recovered tasks approach a factorial centroid geometry. We then analyze a canonical two-view SSL objective and show that its population optimum spans the leading cross-view-stable modes of the associated two-view operator. This yields a closed-form spectral characterization of semantic recoverability: a downstream task is preserved to the extent that its posterior lies in the selected spectral subspace. We validate these predictions on synthetic and real datasets across several SSL methods, testing the predicted relationships among recoverability, directional geometry, spectral structure, and few-shot transfer. Together, these results give a task-level account of what information survives same-instance SSL and how the retained information appears in downstream geometry and transfer.

Figures & tables

Appendix figures & tables7 assets

Supplementary material from the paper’s appendix.

Appendix

Explore similar work

Sep 14, 2026cs.LG

Sharp Rates and a One-Line Correction for Spectral Representation Learning

A self-supervised encoder is trained once, frozen, and reused through lightweight probes on tasks nobody named at training time; the practitioner's question is when the off-the-shelf features are good enough and when they need fixing. Canonical correlation analysis, HGR maximal correlation, and the population optimum of the spectral contrastive loss all return the top-kk singular subspace of a cross-view dependence operator, justified by isotropy: if the task prior has no directional preference, that subspace is universally optimal. We show isotropy is the wrong hypothesis. The prior enters the transfer risk only through the task covariance Λ=E[ΔΔ⊤]Λ=\mathbb{E}[ΔΔ^\top], and only through its compression onto the operator's leading singular directions; what matters is not whether ΛΛ is isotropic but whether its preferred directions are ordered consistently with the operator's spectrum. We prove matching two-sided rates---worst-case regret is exactly 1−1/κ(Λ)1-1/κ(Λ), refines to 1−Ak1-A_k for an alignment coefficient AkA_k, localizes to the top-2k2k subspace, becomes second order under a spectral gap, and is improvable by no task-agnostic representation---and show why alignment is generic: incoherent preferences cancel in high dimension, and TT diverse tasks force α=O~(dx/T)α=\widetilde O(\sqrt{d_x/T}), a quantitative account of why task diversity, not symmetry, makes self-supervised features transfer. The governing statistics cost O(kdx2)O(kd_x^2), and when they signal misalignment a one-line reweighting of the positive-pair term provably restores exact optimality. The result is a diagnostic that answers the practitioner's question from a small labelled budget and refuses when the task bank cannot support the width requested; on controlled data it takes a regret of 0.860.86 down to 0.0030.003, and on a CIFAR-100 encoder it correctly predicts that no correction is needed.
Aug 8, 2026cs.CV

Three Necessary Principles for Self-Supervised Visual Representation Learning

We argue that learning visual representations without labels requires a training signal jointly complete across three non-overlapping objectives: semantic invariance across augmented views, patch-level spatial prediction, and representational non-degeneracy. We formalize these as the observation, prediction, and regularization principles and prove (i) that combining observation and prediction without regularization admits the constant encoder as a global minimizer under negative-free alignment; (ii) that the two objectives are gradient-complementary and structurally non-conflicting at the encoder output; and (iii) that the momentum encoder converges to the same fixed point as the online encoder and provides no collapse guarantee at convergence. Contrastive alignment provides only self-limiting collapse resistance, formalized via an explicit gradient-decay argument. Dropping prediction withholds the spatial training signal by construction; dropping observation forfeits cross-view semantic invariance by construction; at the scale we study, no pair substitutes for the third. Every major self-supervised method is a special case of a single unified energy decomposition. We pair every theoretical claim with a controlled experiment, including a patch-retrieval evaluation for the spatial consequence of prediction.
May 11, 2026cs.CV

Learning to Perceive "Where": Spatial Pretext Tasks for Robust Self-Supervised Learning

Existing self-supervised learning (SSL) methods primarily learn object-invariant representations but often neglect the spatial structure and relationships among object parts. To address this limitation, we introduce Spatial Prediction (SP), a spatially aware pretext regression task that predicts the relative position and scale between a pair of disentangled local views from the same image. By modeling part-to-part relationships in a continuous geometric space, SP encourages representations to capture fine-grained spatial dependencies beyond invariant categorical semantics, thereby learning the compositional structure of visual scenes. SP is implemented as a decoupled plug-in and can be seamlessly integrated into diverse SSL frameworks. Extensive experiments show consistent improvements across image recognition, fine-grained classification, semantic segmentation, and depth estimation, as well as substantial gains in out-of-distribution robustness for object recognition. To evaluate spatial reasoning, we introduce (1) a position and scale prediction task on image patch pairs and (2) a jigsaw understanding task requiring patch reordering and recognition after reconstruction. Strong performance on these tasks indicates improved spatial structure and geometric awareness. Overall, explicitly modeling spatial information provides an effective inductive bias for SSL, leading to more structured representations and better generalization. Code and models will be released.