Bringing Generative Learning to Representation Learning: Self-Supervised Transfer Learning as Distribution Matching
Authors: Yuling Jiao, Wensen Ma, Defeng Sun, Hansheng Wang, Yang Wang
Organizations: School of Artificial Intelligence, Wuhan University, 430072, Wuhan, China · School of Mathematics and Statistics, Wuhan University, Wuhan, 430072, China · Department of Applied Mathematics, The Hong Kong Polytechnic University, Hung Hom, Kowloon, Hong Kong SAR, China · Guanghua School of Management, Peking University, 100871, Beijing, China · Department of Mathematics, The University of Hong Kong, Pokfulam, Hong Kong SAR, China
Most self-supervised learning objectives defend against collapse but leave the target representation law unspecified. We formulate representation learning as Distribution Matching (DM), learning an augmentation-invariant encoder whose induced law matches an explicit geometric reference. The reference law specifies what the learned representation distribution should look like, whereas a separately chosen discrepancy determines how deviations from this target are measured; here we use Mallows distance. The DM framework reveals a directional inverse: generative learning maps a tractable reference to data, whereas representation learning maps data to a designed reference law. We connect the population objective to class-centre separation and classification error and prove a non-asymptotic neural-sieve guarantee. Simulations and image benchmarks show manifold rectification, fine-grained structure and transfer across label spaces.
Figures & tables
Figure 1: Raw Euclidean distance can conflict with semantic meaning. Although x3 is a cropped view of x1 , ∥x1−x2∥2<∥x1−x3∥2 .
Figure 2: The prescribed reference distribution. To draw from PR , a component is selected with probability αk , its orthogonal centre ck is perturbed by the Gaussian direction in ( 4 ), and the result is normalised to the radius- R hypersphere. Two components are shown; ϵ controls their within-component spread.
Figure 3: Error versus unlabelled sample size. The DM error decreases over the four values of nS ; the latent-oracle error is shown as a reference level.
Figure 4: t-SNE visualisation. (a) Standardised observations X ; (b) normalised DM representations f(X) . Colours denote latent classes and are not used in pretraining.
Table 1: Domain-transfer target classification accuracy (%). For each benchmark, the encoder is pretrained on an unlabelled source sample and evaluated on the corresponding target domain using an affine linear probe and a k -NN probe. DM uses K′=384 reference components. The reported DM accuracies are higher than those of the listed baselines in each of the six comparisons.
Feature space
Raw X
Latent oracle Z
DM f(X)
Classification error
0.4160
0.3320
0.3310
Table 4: Classification error after non-linear warping. Here d=512 , nS=15,000 and nT=200 .
Unlabelled size ( nS )
Latent oracle Z
DM f(X)
1,000
0.2970
0.5890
5,000
0.2970
0.4220
10,000
0.2970
0.3065
20,000
0.2970
0.2905
Table 5: Empirical error as nS increases. The same target and test designs are used at every source sample size.
Figure 5: Augmentation-induced semantic bridging. The raw observations x1 and x2 may be far apart, while transformations A1,A2∈A isolate views x1⋆ and x2⋆ whose Euclidean distance is at most δ . The augmentation distance therefore discounts nuisance variation that can be removed by the allowed transformations.
Method
Linear
5 -NN
Seconds per epoch
Optimised DM
92.11
89.17
72.28
DM-Batch
90.22
87.16
≈31
Table 6: CIFAR-10 representation quality and training speed of optimised DM and DM-Batch. Linear and 5 -NN denote test accuracy (%); epoch time is measured on a single Tesla V100.
Explicit geometric references offer a direct way to structure self-supervised representations. Existing adversarial distribution-matching formulations, however, require costly encoder-critic optimization. We introduce Flow-Based Distribution Matching (FBDM), a non-adversarial framework that learns this reference-directed geometry through spherical conditional velocity regression. An ETF-inspired reference allows its number of components K' to exceed the auxiliary flow dimension d* while retaining structured geometric separation. We assign both augmented views of each image to the same target, while limiting how many images each reference center can receive. An explicit alignment loss further pulls the two views' representations closer together. Experiments across benchmarks ranging from CIFAR to ImageNet show that FBDM achieves performance nearly on par with DM and remains competitive with existing SSL methods. Matched training-cost comparisons show a 1.48- to 1.83-fold speedup over DM with a negligible increase in GPU memory usage. We also provide a theoretical explanation for the usefulness of the learned representations: under stated conditions, we bound the downstream misclassification rate in terms of the FBDM pretraining loss.
Yuling Jiao, Wensen Ma, Houduo Qi +1
School of Artificial Intelligence, National Center for Applied Mathematics in Hubei, Hubei Key Laboratory of Computational Science, Wuhan University, Wuhan, China. · Department of Applied Mathematics, The Hong Kong Polytechnic University, Hung Hom, Kowloon, Hong Kong SAR, China. · Department of Data Science and Artificial Intelligence, The Hong Kong Polytechnic University, Hung Hom, Kowloon, Hong Kong SAR, China.
Self-supervised learning (SSL) excels at finding general-purpose latent representations from complex data, yet lacks a unifying theoretical framework that explains the diverse existing methods and guides the design of new ones. We cast SSL as latent distribution matching (LDM): learning representations that maximize their log-probability under an assumed latent model (alignment), while maximizing latent entropy to prevent collapse (uniformity). This view unifies independent component analysis with contrastive, non-contrastive, and predictive SSL methods, including stop gradient approaches. Leveraging LDM, we derive a nonlinear, sampling-free Bayesian filtering model with a Kalman-based predictor for high-dimensional timeseries. We further prove that predictive LDM yields identifiable latent representations under mild assumptions, even with nonlinear predictors. Overall, LDM clarifies the assumptions behind established SSL methods and provides principled guidance for developing new approaches.
Fabian A Mikulasch, Friedemann Zenke
Friedrich Miescher Institute for Biomedical Research, 4056 Basel, Switzerland · Faculty of Science, University of Basel, 4033 Basel, Switzerland
A self-supervised encoder is trained once, frozen, and reused through lightweight probes on tasks nobody named at training time; the practitioner's question is when the off-the-shelf features are good enough and when they need fixing. Canonical correlation analysis, HGR maximal correlation, and the population optimum of the spectral contrastive loss all return the top-k singular subspace of a cross-view dependence operator, justified by isotropy: if the task prior has no directional preference, that subspace is universally optimal. We show isotropy is the wrong hypothesis. The prior enters the transfer risk only through the task covariance Λ=E[ΔΔ⊤], and only through its compression onto the operator's leading singular directions; what matters is not whether Λ is isotropic but whether its preferred directions are ordered consistently with the operator's spectrum. We prove matching two-sided rates---worst-case regret is exactly 1−1/κ(Λ), refines to 1−Ak for an alignment coefficient Ak, localizes to the top-2k subspace, becomes second order under a spectral gap, and is improvable by no task-agnostic representation---and show why alignment is generic: incoherent preferences cancel in high dimension, and T diverse tasks force α=O(dx/T), a quantitative account of why task diversity, not symmetry, makes self-supervised features transfer. The governing statistics cost O(kdx2), and when they signal misalignment a one-line reweighting of the positive-pair term provably restores exact optimality. The result is a diagnostic that answers the practitioner's question from a small labelled budget and refuses when the task bank cannot support the width requested; on controlled data it takes a regret of 0.86 down to 0.003, and on a CIFAR-100 encoder it correctly predicts that no correction is needed.
Dier Tang, Jing Yee Tan, Guangyue Han
Department of Mathematics, The University of Hong Kong