Learning a Flow to Self-Supervised Representations
Authors: Yuling Jiao, Wensen Ma, Houduo Qi, Defeng Sun
Organizations: School of Artificial Intelligence, National Center for Applied Mathematics in Hubei, Hubei Key Laboratory of Computational Science, Wuhan University, Wuhan, China. · Department of Applied Mathematics, The Hong Kong Polytechnic University, Hung Hom, Kowloon, Hong Kong SAR, China. · Department of Data Science and Artificial Intelligence, The Hong Kong Polytechnic University, Hung Hom, Kowloon, Hong Kong SAR, China.
Explicit geometric references offer a direct way to structure self-supervised representations. Existing adversarial distribution-matching formulations, however, require costly encoder-critic optimization. We introduce Flow-Based Distribution Matching (FBDM), a non-adversarial framework that learns this reference-directed geometry through spherical conditional velocity regression. An ETF-inspired reference allows its number of components K' to exceed the auxiliary flow dimension d* while retaining structured geometric separation. We assign both augmented views of each image to the same target, while limiting how many images each reference center can receive. An explicit alignment loss further pulls the two views' representations closer together. Experiments across benchmarks ranging from CIFAR to ImageNet show that FBDM achieves performance nearly on par with DM and remains competitive with existing SSL methods. Matched training-cost comparisons show a 1.48- to 1.83-fold speedup over DM with a negligible increase in GPU memory usage. We also provide a theoretical explanation for the usefulness of the learned representations: under stated conditions, we bound the downstream misclassification rate in terms of the FBDM pretraining loss.
Figures & tables
CIFAR-10
CIFAR-100
STL-10
Tiny ImageNet
Method
Linear
5 -NN
Linear
5 -NN
Linear
5 -NN
Linear
5 -NN
FBDM
92.37
89.68
66.59
56.74
89.79
86.23
48.55
33.01
DM ( Jiao et al., 2026 )
92.11
89.17
67.71
56.18
90.22
85.51
–
Barlow Twins ( Zbontar et al., 2021 )
88.51
86.53
65.78
55.76
88.36
83.71
47.44
32.65
SimCLR ( Chen et al., 2020 )
91.80
88.42
66.83
56.56
90.51
85.68
48.84
32.86
BYOL ( Grill et al., 2020 )
91.73
89.45
66.60
56.82
91.99
88.64
51.00
36.24
Table 1: Top-1 accuracy (%) on CIFAR-10, CIFAR-100, STL-10, and Tiny ImageNet with ResNet-18. Linear evaluates the frozen representation; 5 -NN uses k=5 . Non-FBDM results except DM follow the published benchmark ( Weng et al., 2022 ) ; DM results are taken from its source paper ( Jiao et al., 2026 ) .
Method
Linear
FBDM
59.32
SimCLR ( Chen et al., 2020 )
60.14
Table 2: ImageNet-1K linear-evaluation top-1 accuracy (%) with ResNet-50 after 100 epochs at global batch size 512. The SimCLR value is the reproduction reported by AndrewAtanov/simclr-pytorch .
Dataset
System
Peak memory (GiB)
Seconds/epoch
V100-h/1000 ep.
Speedup vs. DM
CIFAR-10
DM
4.78/5.14
75.61
21.00
1.83×
FBDM
5.28/5.82
41.35
11.49
CIFAR-100
DM
4.78/5.14
76.37
21.21
1.69×
FBDM
5.28/5.82
45.27
12.58
STL-10
DM
7.06/8.17
173.35
48.15
1.48×
FBDM
7.56/8.94
117.39
32.61
Table 3: Native-stack training cost on one Tesla V100-SXM2-16GB. Memory is peak CUDA allocated/reserved GiB. V100-hours are projected from the median seconds per epoch for 1000 epochs.
Appendix figures & tables6 assets
Supplementary material from the paper’s appendix.
Appendix
Symbol
Meaning / definition
First defined
Z0
Encoder representation f(X) before ODE evolution; used for downstream classification.
Section 3.3
R
Assigned reference target; the prescribed endpoint of the analytic path.
Section 3.3
Zt
State on the prescribed analytic path, with Z1=R .
( 10 )
Ut
Conditional target velocity dZt/dt .
( 11 )
Φt
Learned ODE map satisfying dΦt(z)/dt=vϕ(Φt(z),t) and Φ0(z)=z .
( 25 )
Zt
Learned ODE state Φt(Z0) ; Z1 is its terminal representation, not the prescribed target.
( 25 )
Appendix
Table 4: Key notation for the flow and classification analysis. The index k denotes a semantic class, and j denotes a reference component. All representation vectors lie in Rd⋆ .
Dataset
B
K′
d⋆
mDN(e)
CIFAR-10
512
128
64
5(e<100),7(100≤e<250),9(e≥250)
CIFAR-100
512
160
64
5(e<125),7(125≤e<275),9(e≥275)
STL-10
512
128
64
5(e<100),7(100≤e<250),9(e≥250)
Tiny ImageNet
1024
320
96
5(e<100),7(100≤e<250),9(e≥250)
ImageNet-1K
512
4096
512
4
Appendix
Table 5: Dataset-specific assignment capacities used by the reported FBDM runs. Here B is the minibatch size and mDN(e) is the number of available slots per center at epoch e .
Dataset
Reported entries
Clock
Sampling convention
CIFAR-10
Linear, 5 -NN
rational, a=0.5
path-uniform
CIFAR-100
Linear, 5 -NN
rational, a=0.5
path-uniform
STL-10
Linear, 5 -NN
exponential, a=1
path-uniform
Tiny ImageNet
Linear, 5 -NN
exponential, a=2
path-uniform
ImageNet-1K
Linear
rational, a=0.5
path-uniform
Appendix
Table 6: Time reparameterization for the FBDM entries in Tables 1 and 2 . Path-uniform sampling draws S∼Unif([0,1]) and sets T=s−1(S) .
Component
Setting
Backbone
ResNet-18
Reference geometry
Simplex-spectral
K′ , d⋆ , and capacity
See Table 5
Center perturbation
10−3
Velocity time feature
raw t
Flow path
Spherical geodesic
Appendix
Table 7: Core configuration for the reported ResNet-18 FBDM results on CIFAR-10, CIFAR-100, STL-10, and Tiny ImageNet.
Spherical path
Linear path
Dataset
Linear
5 -NN
Linear
5 -NN
CIFAR-10
92.37
89.68
90.50
88.47
CIFAR-100
66.59
56.74
65.01
55.49
STL-10
89.79
86.23
87.94
83.85
Tiny ImageNet
48.55
33.01
47.47
32.65
Appendix
Table 8: Reported FBDM top-1 accuracy (%) with the spherical path used in Table 1 and with direct linear interpolation. Linear denotes frozen-representation linear evaluation; 5 -NN uses k=5 .
Factor
Variant / control
Δ Linear
Δ 5-NN
FM encoder gradient
0.75/1.00
+0.860
+0.57
LR warmup steps
0/500
+0.330
+0.56
View blur (p1,p2)
(.20,.30)/(.25,.25)
+0.420
+0.08
Time-loss exponent
0.5/0
+0.033
+0.89
Appendix
Table 9: Strict fixed-1000 paired ablations on CIFAR-10. Each row compares a variant with its matched control; Δ is variant minus control.
Most self-supervised learning objectives defend against collapse but leave the target representation law unspecified. We formulate representation learning as Distribution Matching (DM), learning an augmentation-invariant encoder whose induced law matches an explicit geometric reference. The reference law specifies what the learned representation distribution should look like, whereas a separately chosen discrepancy determines how deviations from this target are measured; here we use Mallows distance. The DM framework reveals a directional inverse: generative learning maps a tractable reference to data, whereas representation learning maps data to a designed reference law. We connect the population objective to class-centre separation and classification error and prove a non-asymptotic neural-sieve guarantee. Simulations and image benchmarks show manifold rectification, fine-grained structure and transfer across label spaces.
Yuling Jiao, Wensen Ma, Defeng Sun +2
School of Artificial Intelligence, Wuhan University, 430072, Wuhan, China · School of Mathematics and Statistics, Wuhan University, Wuhan, 430072, China · Department of Applied Mathematics, The Hong Kong Polytechnic University, Hung Hom, Kowloon, Hong Kong SAR, China +2
Flow matching (FM) trains a time-dependent vector field that transports samples from a simple prior to a complex data distribution. However, for high-dimensional images, each training sample supervises only a single trajectory and intermediate point, yielding an extremely sparse and high-variance training signal. This under-constrained supervision can cause flow collapse, where the learned dynamics memorize specific source-target pairings, mapping diverse inputs to overly similar outputs, failing to generalize. We introduce Posterior-Augmented Flow Matching (PAFM), a theoretically grounded generalization of FM that replaces single-target supervision with an expectation over an approximate posterior of valid target completions for a given intermediate state and condition. PAFM factorizes this intractable posterior into (i) the likelihood of the intermediate under a hypothesized endpoint and (ii) the prior probability of that endpoint under the condition, and uses an importance sampling scheme to construct a mixture over multiple candidate targets. We prove that PAFM yields an unbiased estimator of the original FM objective while substantially reducing gradient variance during training by aggregating information from many plausible continuation trajectories per intermediate. Finally, we show that PAFM improves over FM by up to 3.4 FID50K across different model scales (SiT-B/2 and SiT-XL/2), different architectures (SiT and MMDiT), and in both class and text conditioned benchmarks (ImageNet and CC12M), with a negligible increase in the compute overhead. Code: https://github.com/gstoica27/PAFM.git.
George Stoica, Sayak Paul, Matthew Wallingford +6
Georgia Tech · University of Washington · Hugging Face +2
Dataset distillation seeks to synthesize a compact surrogate dataset that enables performance comparable to training on the original dataset for downstream tasks. For the scenario where pre-trained self-supervised models serve as priors, traditional Linear Gradient Matching optimizes synthetic images by encouraging them to mimic the gradient updates induced by real images on the linear probe. However, this batch-level formulation requires loading thousands of real images and applying multiple differentiable augmentations to synthetic images at each distillation step, leading to substantial computational and memory overheads. In this paper, we revisit the linear gradient and theoretically derive that it is essentially a local relative distribution directed from target class centers toward non-target class centers, which we term flow. This property causes suboptimality and instability, often necessitating expensive multiple augmentations to compensate. To address this, we introduce Statistical Flow Matching, an optimal, stable, and efficient supervised learning framework that optimizes synthetic images by aligning global statistical flows in the original data. Our approach loads raw statistics only once and performs a single augmentation pass on the synthetic data, achieving performance comparable to or better than the state-of-the-art method with 10x less GPU memory usage and 4x faster distillation time. Moreover, increasing the number of augmentations for our method yields further performance gains while incurring lower additional cost. Our code is publicly available at https://github.com/einsteinxia/SFM.
Qianxin Xia, Jiawei Du, Yuhan Zhang +6
University of Electronic Science and Technology of China, Chengdu, China · CFAR, Singapore · SWJTU, Chengdu, China