Dataset distillation seeks to synthesize a compact surrogate dataset that enables performance comparable to training on the original dataset for downstream tasks. For the scenario where pre-trained self-supervised models serve as priors, traditional Linear Gradient Matching optimizes synthetic images by encouraging them to mimic the gradient updates induced by real images on the linear probe. However, this batch-level formulation requires loading thousands of real images and applying multiple differentiable augmentations to synthetic images at each distillation step, leading to substantial computational and memory overheads. In this paper, we revisit the linear gradient and theoretically derive that it is essentially a local relative distribution directed from target class centers toward non-target class centers, which we term flow. This property causes suboptimality and instability, often necessitating expensive multiple augmentations to compensate. To address this, we introduce Statistical Flow Matching, an optimal, stable, and efficient supervised learning framework that optimizes synthetic images by aligning global statistical flows in the original data. Our approach loads raw statistics only once and performs a single augmentation pass on the synthetic data, achieving performance comparable to or better than the state-of-the-art method with 10x less GPU memory usage and 4x faster distillation time. Moreover, increasing the number of augmentations for our method yields further performance gains while incurring lower additional cost. Our code is publicly available at https://github.com/einsteinxia/SFM.
Figures & tables
Figure 1 : Comparison between LGM and our SFM in terms of distillation time, GPU memory usage, and validation accuracy on ImageNet-100 with IPC=1. The size of the circle denotes different amounts of Augmentations Per Batch (APB). The distillation model is EVA-02 and the generalized models are CLIP, DINO-v2 and MoCo-v3. With only one augmentation, our SFM beats LGM using multiple augmentations, which substantially reduces resource consumption while achieving state-of-the-art validation performance.
Figure 2 : Comparison of the pipelines: Linear Gradient Matching vs. Statistical Flow Matching. During the original data loading phase, LGM dynamically loads a large batch of samples at every distillation step per epoch, whereas SFM computes the statistics only once and then fixes them during subsequent optimization. During the synthetic data distillation phase, LGM employs multiple augmentations to effectively align gradients or capture local relative distributions, while retaining expensive model computation graphs for backpropagation. In contrast, SFM requires only a single augmentation to stably and efficiently match the optimal statistical flow of the original data.
Mode
CLIP
DINO-v2
EVA-02
MoCo-v3
Random
64.9 ± 0.2
78.9 ± 0.1
72.2 ± 0.1
72.5 ± 0.1
Fixed
64.4 ± 0.4
78.4 ± 0.2
72.6 ± 0.3
72.4 ± 0.2
Analytic
64.3 ± 0.4
78.4 ± 0.2
73.0 ± 0.1
71.9 ± 0.2
Table 1: A comparison of average generalization performance under different W settings. The number of augmentation per batch is set to 1. It proves to be insensitive to changes in W .
Train Set (1 Img/Cls)
ImageNet-100
ImageNet-1k
CLIP
DINO-v2
EVA-02
MoCo-v3
Average
CLIP
DINO-v2
EVA-02
MoCo-v3
Average
Random
56.6 ± 1.6
74.8 ± 2.8
64.5 ± 2.7
61.4 ± 2.6
64.3 ± 2.4
31.7 ± 0.5
50.3 ± 0.5
37.7 ± 0.4
38.8 ± 0.6
39.6 ± 0.5
Centroids
77.1 ± 0.1
86.9 ± 0.3
80.9 ± 0.2
77.7 ± 0.1
80.6 ± 0.2
53.9 ± 0.0
69.5 ± 0.1
58.1 ± 0.1
57.4 ± 0.0
59.7 ± 0.1
Neighbors (LGM*)
67.8 ± 0.3
86.0 ± 0.2
78.8 ± 0.2
77.1 ± 0.1
77.4 ± 0.2
38.8 ± 0.1
67.7 ± 0.1
49.9 ± 0.1
56.4 ± 0.0
53.2 ± 0.1
Neighbors (SFM)
72.2 ± 0.2
88.0 ± 0.1
80.4 ± 0.0
77.8 ± 0.1
79.6 ± 0.1
48.1 ± 0.1
69.6 ± 0.0
57.3 ± 0.0
57.0 ± 0.0
58.0 ± 0.0
LGM
77.3 ± 0.2
89.8 ± 0.0
85.7 ± 0.1
81.1 ± 0.1
83.5 ± 0.1
55.5 ± 0.0
73.5 ± 0.2
67.6 ± 0.0
61.6 ± 0.0
64.6 ± 0.1
Table 2 : The distillation performance on specific models. We compare our SFM with LGM and several real-image baselines on various datasets. Our SFM, along with its derivatives SFM* and Neighbors (SFM), consistently outperforms other relevant baselines. The superscript “” indicates multiple augmentations, otherwise a single augmentation. “” indicates out-of-memory due to resource limitations. The best and second-best results are highlighted in bold and underlined .
Distill. Model
CLIP
DINO-v2
Eval. Model
CLIP
DINO-v2
EVA-02
MoCo-v3
Average
CLIP
DINO-v2
EVA-02
MoCo-v3
Average
Centroids
77.1 ± 0.1
85.1 ± 0.2
82.6 ± 0.2
74.7 ± 0.1
79.9 ± 0.2
72.4 ± 0.0
86.9 ± 0.3
80.1 ± 0.4
75.8 ± 0.1
78.8 ± 0.2
Neighbors (LGM*)
67.8 ± 0.3
81.1 ± 0.1
75.0 ± 0.2
70.4 ± 0.0
73.5 ± 0.1
71.2 ± 0.1
86.0 ± 0.2
78.0 ± 0.4
74.7 ± 0.0
77.6 ± 0.2
Neighbors (SFM)
72.2 ± 0.2
83.6 ± 0.2
75.8 ± 0.0
72.3 ± 0.1
76.0 ± 0.1
72.7 ± 0.0
88.0 ± 0.1
80.0 ± 0.2
75.4 ± 0.0
79.0 ± 0.1
LGM
77.3 ± 0.2
70.2 ± 0.4
70.4 ± 0.1
41.8 ± 0.2
64.9 ± 0.2
70.8 ± 0.1
89.8 ± 0.0
82.4 ± 0.2
72.5 ± 0.1
78.9 ± 0.1
SFM (Ours)
86.7 ± 0.1
84.9 ± 0.3
85.5 ± 0.2
68.7 ± 0.0
81.5 ± 0.2
78.6 ± 0.0
90.6 ± 0.0
86.2 ± 0.1
80.5 ± 0.0
84.0 ± 0.0
Table 3 : The generalization performance across models. We synthesize images using a given model and then evaluate them across other models on ImageNet-100. Compared to other baselines, our distilled datasets generalize well and achieve optimal performance across different model pairs, aside from outlier pairs of CLIP and MoCo-v3, which is caused by discrepancies or inconsistencies inherent in the models themselves. The superscript “*” indicates multiple augmentations, otherwise a single augmentation. The best and second-best results are highlighted in bold and underlined .
Figure 6Table 7
Figure 4 : A comparison between soft label (solid line) with different temperatures and our CI (dashed line). Our CI outperforms soft labels across all settings.
Figure 5 : A comparison of images synthesized by LGM versus our SFM under a single augmentation, alongside LGM* and SFM* under multiple augmentations. The dataset is Stanford Dogs.
Figure 6 : The (a) distribution, along with the flows of (b) LGM and (c) our SFM, on ImageNet-100. We perform PCA for dimensionality reduction, visualizing 20 classes. Both LGM and our SFM synthesize samples on the distribution boundary with discriminative attributes.
Appendix figures & tables9 assets
Supplementary material from the paper’s appendix.
Appendix
Mean
CLIP
DINO-v2
EVA-02
MoCo-v3
0.5
81.0 ± 0.2
88.0 ± 0.2
88.9 ± 0.2
70.5 ± 0.0
0
82.2 ± 0.0
88.8 ± 0.0
88.6 ± 0.3
76.6 ± 0.0
Appendix
Table 7: Effect of different gaussian noise means on performance. The distillation model is EVA-02.
Figure 7 : An overview of our Classifier Inheritance (CI). We first train a classifier on the original dataset using a frozen distillation model, then freeze it. During evaluation on synthetic images, we directly drive the evaluation model to align its feature representation with that of the distillation model, and use the aforementioned classifier for inference and prediction.
Epoch
ImageNet-100
ImageNet-1k
CLIP
DINO-v2
EVA-02
MoCo-v3
Average
CLIP
DINO-v2
EVA-02
MoCo-v3
Average
10
92.8 ± 0.1
95.3 ± 0.1
94.5 ± 0.0
88.2 ± 0.1
92.7 ± 0.1
79.0 ± 0.0
82.0 ± 0.0
82.4 ± 0.1
73.9 ± 0.0
79.3 ± 0.0
Epoch
Stanford Dogs
CUB-200
CLIP
DINO-v2
EVA-02
MoCo-v3
Average
CLIP
DINO-v2
EVA-02
MoCo-v3
Average
10
72.0 ± 0.1
88.8 ± 0.0
83.2 ± 0.1
70.4 ± 0.3
78.6 ± 0.1
72.5 ± 0.4
89.7 ± 0.1
83.2 ± 0.1
45.3 ± 0.4
72.7 ± 0.3
Appendix
Table 8: The performance of the gold classifier trained on various full datasets.
Distill. Model
CLIP
DINO-v2
Eval. Model
CLIP
DINO-v2
EVA-02
MoCo-v3
Average
CLIP
DINO-v2
EVA-02
MoCo-v3
Average
LGM*
63.0 ± 0.0
56.4 ± 0.1
59.7 ± 0.1
39.5 ± 0.0
54.7 ± 0.1
54.1 ± 0.0
75.0 ± 0.1
65.4 ± 0.1
60.0 ± 0.0
63.7 ± 0.1
LGM
55.5 ± 0.0
45.0 ± 0.0
48.5 ± 0.0
25.8 ± 0.0
43.7 ± 0.0
46.7 ± 0.0
73.5 ± 0.2
61.2 ± 0.1
54.7 ± 0.1
59.0 ± 0.1
SFM
67.1 ± 0.0
64.2 ± 0.1
64.7 ± 0.0
46.6 ± 0.0
60.7 ± 0.1
56.6 ± 0.0
75.1 ± 0.1
65.6 ± 0.0
59.9 ± 0.0
64.3 ± 0.0
Distill. Model
EVA-02
MoCo-v3
Eval. Model
CLIP
DINO-v2
EVA-02
MoCo-v3
Average
CLIP
DINO-v2
EVA-02
MoCo-v3
Average
Appendix
Table 9 : The cross-model generalization performance on ImageNet-1k.
Distill. Model
CLIP
DINO-v2
Eval. Model
CLIP
DINO-v2
EVA-02
MoCo-v3
Average
CLIP
DINO-v2
EVA-02
MoCo-v3
Average
LGM
16.4 ± 0.0
15.2 ± 0.0
13.6 ± 0.0
2.7 ± 0.1
12.0 ± 0.0
22.1 ± 0.2
81.1 ± 0.1
50.3 ± 0.1
26.0 ± 0.1
44.9 ± 0.1
SFM
67.6 ± 0.2
51.4 ± 0.3
48.3 ± 0.3
16.2 ± 0.3
45.9 ± 0.3
37.8 ± 0.1
85.7 ± 0.1
61.7 ± 0.1
29.5 ± 0.1
53.7 ± 0.1
Distill. Model
EVA-02
MoCo-v3
Eval. Model
CLIP
DINO-v2
EVA-02
MoCo-v3
Average
CLIP
DINO-v2
EVA-02
MoCo-v3
Average
LGM
17.8 ± 0.1
37.4 ± 0.1
51.5 ± 0.2
8.9 ± 0.1
28.9 ± 0.1
6.4 ± 0.0
36.4 ± 0.5
17.5 ± 0.3
25.5 ± 0.0
21.5 ± 0.2
Appendix
Table 10 : The cross-model generalization performance on CUB-200.
Distill. Model
CLIP
DINO-v2
Eval. Model
CLIP
DINO-v2
EVA-02
MoCo-v3
Average
CLIP
DINO-v2
EVA-02
MoCo-v3
Average
LGM
15.0 ± 0.0
17.7 ± 0.2
16.3 ± 0.0
6.2 ± 0.1
13.8 ± 0.1
27.5 ± 0.1
79.8 ± 0.2
59.7 ± 0.0
56.7 ± 0.6
55.9 ± 0.2
SFM
55.0 ± 0.3
62.6 ± 0.7
60.0 ± 0.3
45.6 ± 0.8
55.8 ± 0.5
39.2 ± 0.2
81.8 ± 0.2
66.0 ± 0.4
62.3 ± 0.2
62.3 ± 0.3
Distill. Model
EVA-02
MoCo-v3
Eval. Model
CLIP
DINO-v2
EVA-02
MoCo-v3
Average
CLIP
DINO-v2
EVA-02
MoCo-v3
Average
LGM
16.0 ± 0.0
53.8 ± 0.3
65.0 ± 0.2
32.5 ± 0.2
41.8 ± 0.2
11.3 ± 0.3
66.1 ± 0.4
48.9 ± 0.1
63.5 ± 0.2
47.5 ± 0.3
Appendix
Table 11 : The cross-model generalization performance on Stanford Dogs.
Figure 8 : Distilling the ArtBench Liao et al. (2022) dataset.
Size
ViT-S(dim=384)
ViT-B(dim=768)
ViT-L(dim=1024)
SFM
84.9 ± 0.3
90.6 ± 0.0
90.4 ± 0.3
SFM+CI
86.0 ± 0.1
95.1 ± 0.1
92.4 ± 0.1
Appendix
Table 12: Performance of our SFM and CI under homogeneous architectures with different scales.
Augmentation
1
3
5
7
10
LGM
GPU(GB)
16.6
48.5
84.5
107.3
165.2
SFM
16.3
48.1
84.2
106.8
164.5
LGM
Time(min)
21
35
48
63
81
SFM
18
28
38
48
62
Appendix
Table 13: Quantitative comparison between LGM and our SFM in terms of distillation time and GPU memory usage. The distillation model is EVA-02.
Dataset distillation (DD) condenses a large original dataset into a small distilled dataset with high training utility. Decoupled statistical matching methods substantially reduce distillation time and memory overhead while achieving strong performance. However, they typically supervise all distilled samples using running statistics estimated from the entire original dataset. These statistics mainly capture the average feature distribution while overlooking differences in sample difficulty, limiting their ability to characterize the difficulty structure of the original data. To address this issue, we propose Precise Statistical Matching (PSM) by difficulty. After pretraining, PSM uses the Global Precision Score (GPS) to estimate image difficulty, ranks the samples within each class, and partitions each class into IPC (images per class) difficulty groups. During distillation, Statistics Updated Again (SUA) updates the teacher's batch normalization (BN) running statistics through forward passes on original samples from each group, providing difficulty-specific supervision for the corresponding distilled batch. Meanwhile, Initial Sample Screening (ISS) initializes distilled samples using original images from the corresponding difficulty group, providing an effective starting point for precise matching. Experiments across multiple datasets and model architectures demonstrate that PSM broadens the difficulty range of distilled samples and improves downstream performance in most evaluated settings. Code will be released.
Hongxu Ma, Guang Li, Shijie Wang +6
Zhejiang University · Hokkaido University · The University of Queensland +3
We revisit dataset distillation from an outcome-centric perspective. Rather than aligning process surrogates (per-step gradients or training trajectories), Influence Matching (Inf-Match) aligns the final outcome of training: it learns a compact synthetic set whose effect on the converged parameters matches that of the full dataset. Concretely, we introduce a fully differentiable, sample-level influence estimator that quantifies parameter shifts from adding or removing data, without time-consuming inverse-Hessian products or convexity assumptions. The estimator runs in linear time by unrolling the optimization dynamics and applying a first-order Taylor approximation. We then learn the synthetic set by minimizing the mismatch between its influence and that of the real dataset, yielding outcome alignment rather than heuristic process imitation. Inf-Match delivers the best accuracy across standard classification benchmarks. For instance, on Tiny-ImageNet (IPC=10), Inf-Match attains 31.5%, a +4.7% improvement over NCFM. Beyond classification, Inf-Match scales to vision-language distillation on Flickr30K, outperforming strong process-matching baselines. For instance, with 200 to 1000 synthetic samples, our method achieved a leading impressive average on image/text retrieval tasks, higher than NCFM by 2.5%. The code will be released via https://github.com/hrtan/infmatch.
Explicit geometric references offer a direct way to structure self-supervised representations. Existing adversarial distribution-matching formulations, however, require costly encoder-critic optimization. We introduce Flow-Based Distribution Matching (FBDM), a non-adversarial framework that learns this reference-directed geometry through spherical conditional velocity regression. An ETF-inspired reference allows its number of components K' to exceed the auxiliary flow dimension d* while retaining structured geometric separation. We assign both augmented views of each image to the same target, while limiting how many images each reference center can receive. An explicit alignment loss further pulls the two views' representations closer together. Experiments across benchmarks ranging from CIFAR to ImageNet show that FBDM achieves performance nearly on par with DM and remains competitive with existing SSL methods. Matched training-cost comparisons show a 1.48- to 1.83-fold speedup over DM with a negligible increase in GPU memory usage. We also provide a theoretical explanation for the usefulness of the learned representations: under stated conditions, we bound the downstream misclassification rate in terms of the FBDM pretraining loss.
Yuling Jiao, Wensen Ma, Houduo Qi +1
School of Artificial Intelligence, National Center for Applied Mathematics in Hubei, Hubei Key Laboratory of Computational Science, Wuhan University, Wuhan, China. · Department of Applied Mathematics, The Hong Kong Polytechnic University, Hung Hom, Kowloon, Hong Kong SAR, China. · Department of Data Science and Artificial Intelligence, The Hong Kong Polytechnic University, Hung Hom, Kowloon, Hong Kong SAR, China.