Diffusion-based dataset distillation (DD) suffers from a fundamental objective mismatch: likelihood-driven diffusion models prioritize density approximation over the discriminative decision boundaries required for downstream tasks. Beyond semantic mismatch, relying solely on density also leads to geometric coverage loss, where generated samples collapse into a few high-density modes and fail to cover the manifold's structural diversity. We propose Manifold-Guided Policy Optimization (MGPO), which reformulates DD as a multi-objective reinforcement learning problem and achieves Dual-Space Alignment via a pixel-space discriminative reward and a latent-space geometric reward guided by a class-wise Minimum Spanning Tree (MST). The discriminative reward enforces class separability, while the MST-based geometric reward encourages generated latents to cover a sparse geometric skeleton of each class, jointly addressing both failure modes. We further provide an idealized analysis that motivates the MST-based reward, including a Hausdorff approximation bound and a subsampling bound independent of the dataset size. The reward-modular design extends to structured tasks such as object detection and segmentation by substituting the frozen task reward model. Extensive experiments show MGPO consistently outperforms existing methods, including a +8.0% mIoU gain on segmentation under low-budget settings.
Figures & tables
Figure 1: t-SNE of latent distributions on ImageWoof. Blue density: real data; orange dots: Minimax Diffusion; red dots: MGPO (ours); red stars with green edges: class-wise MST skeleton. MGPO covers a broader area along the MST skeleton, indicating improved manifold coverage.
Figure 2: Overview of the proposed MGPO framework. (a) Pixel Reward ( rdisc ) is computed from downstream objectives such as classification, detection, and segmentation, encouraging semantic correctness and task consistency. (b) MST Reward ( rgeo ) is constructed in latent space. Real dataset features are encoded by encoder and projected via PCA to obtain a structured latent representation. A kNN graph is built to approximate local neighborhood structure, followed by a Minimum Spanning Tree (MST) to capture the global manifold topology. The generated latent samples are aligned with this manifold structure, promoting both representativeness and structural consistency.
IPC (Ratio)
Model
Random
K-Center
Herding
DiT
Minimax
DM
IDC-1
D 4 M
DDVLCP
ManifoldGD
MGPO (Ours)
Full
10 (0.8%)
ConvNet-6
24.3 ± 1.1
19.4 ± 0.9
26.7 ± 0.5
34.2 ± 1.1
37.0 ± 1.0
26.9 ± 1.2
33.3 ± 1.1
29.4 ± 0.9
34.8 ± 2.4
36.9 ± 0.6
37.2 ± 0.8
86.4 ± 0.2
ResNetAP-10
29.4 ± 0.8
22.1 ± 0.1
32.0 ± 0.3
34.7 ± 0.5
39.2 ± 1.3
30.3 ± 1.2
39.1 ± 0.5
33.2 ± 2.1
39.5 ± 1.5
38.3 ± 0.4
40.9 ± 1.1
87.5 ± 0.5
ResNet-18
27.7 ± 0.9
21.1 ± 0.4
30.2 ± 1.2
34.7 ± 0.4
37.6 ± 0.9
33.4 ± 0.7
37.3 ± 0.2
32.3 ± 1.2
39.9 ± 2.6
39.2 ± 0.7
40.1 ± 0.3
89.3 ± 1.2
20 (1.6%)
ConvNet-6
29.1 ± 0.7
21.5 ± 0.8
29.5 ± 0.3
36.1 ± 0.8
37.6 ± 0.2
29.9 ± 1.0
35.5 ± 0.8
34.0 ± 2.3
37.9 ± 1.9
37.7 ± 0.6
39.6 ± 1.3
86.4 ± 0.2
ResNetAP-10
32.7 ± 0.4
25.1 ± 0.7
34.9 ± 0.1
41.1 ± 0.8
45.8 ± 0.5
35.2 ± 0.6
43.4 ± 0.3
40.1 ± 1.6
44.5 ± 2.2
45.6 ± 0.5
46.1 ± 0.3
87.5 ± 0.5
ResNet-18
29.7 ± 0.5
23.6 ± 0.3
32.2 ± 0.6
40.5 ± 0.5
42.5 ± 0.6
29.8 ± 1.7
38.6 ± 0.2
38.4 ± 1.1
44.5 ± 2.0
42.4 ± 0.4
45.9 ± 0.2
89.3 ± 1.2
Table 1: Comparison of SOTA dataset distillation methods on ImageWoof under various IPC and backbone settings. The best results are marked as bold, and the second are underlined.
IPC
Random
DiT
Minimax
D 4 M
DDVLCP
MGPO (Ours)
Nette
10
54.2 ± 1.6
59.1 ± 0.7
58.6 ± 0.4
60.9 ± 1.7
64.8 ± 3.6
66.9 ± 2.2
20
63.5 ± 0.5
64.8 ± 1.2
70.6 ± 0.3
66.3 ± 1.3
71.4 ± 0.5
73.9 ± 0.7
50
76.1 ± 1.1
73.3 ± 0.9
83.7 ± 0.4
77.7 ± 1.1
81.2 ± 0.8
85.4 ± 1.1
IDC
10
48.1 ± 0.8
54.1 ± 0.4
51.9 ± 1.4
52.8 ± 0.5
57.0 ± 1.4
54.6 ± 1.7
20
52.5 ± 0.9
58.9 ± 0.2
59.1 ± 3.7
58.5 ± 0.4
63.3 ± 1.2
64.8 ± 0.4
50
68.1 ± 0.7
64.3 ± 0.6
69.4 ± 1.4
69.1 ± 0.8
71.9 ± 0.4
77.8 ± 0.9
Table 2: Comparison on ImageNette and ImageIDC under different IPC settings. Results are obtained on ResNet-18. The best results are marked as bold and the second are underlined.
IPC
SRe 2 L
RDED
DiT
Minimax
MGPO (Ours)
10
21.3 ± 0.6
42.0 ± 0.1
39.6 ± 0.4
44.3 ± 0.5
45.1 ± 0.6
50
46.8 ± 0.2
56.5 ± 0.1
52.9 ± 0.6
58.6 ± 0.3
59.1 ± 0.4
Table 3: Comparison of dataset distillation methods on ImageNet-1K under different IPC settings.
Method
Pascal VOC
MS COCO
0.5%
1%
2%
0.25%
0.5%
1%
mAP
AP 50
mAP
AP 50
mAP
AP 50
mAP
AP 50
mAP
AP 50
mAP
AP 50
Random
0.8 ± 0.2
3.1 ± 0.4
4.2 ± 0.5
13.7 ± 0.6
12.4 ± 0.4
34.3 ± 0.5
0.5 ± 0.1
1.7 ± 0.3
3.7 ± 0.2
10.1 ± 0.3
7.2 ± 0.8
17.3 ± 0.9
Uniform
0.9 ± 0.1
3.4 ± 0.3
5.7 ± 0.2
17.7 ± 0.4
13.8 ± 0.3
36.2 ± 0.4
0.8 ± 0.2
2.4 ± 0.5
3.4 ± 0.4
9.5 ± 0.6
7.4 ± 0.5
17.6 ± 0.5
K-Center
0.5 ± 0.1
2.1 ± 0.3
3.6 ± 0.6
12.3 ± 0.3
10.9 ± 0.6
29.3 ± 0.6
0.4 ± 0.1
1.5 ± 0.2
3.2 ± 0.5
9.5 ± 0.5
6.1 ± 0.3
15.4 ± 0.6
Herding
0.6 ± 0.2
2.4 ± 0.2
3.5 ± 0.5
11.9 ± 0.5
10.4 ± 0.4
28.7 ± 0.7
0.5 ± 0.1
1.8 ± 0.4
3.5 ± 0.3
9.7 ± 0.3
6.7 ± 0.4
16.3 ± 0.7
Table 4: Performance comparison with coreset selection methods. Metrics are reported as mAP and AP 50 (%) across different selection ratios.
Table 7Figure 8
Appendix figures & tables7 assets
Supplementary material from the paper’s appendix.
Appendix
Backbone
IPC
SRe 2 L ( Yin et al., 2023 )
Minimax ( Gu et al., 2024 )
RDED ( Sun et al., 2024 )
CaO 2 ( Wang et al., 2025 )
MGPO (Ours)
ResNet-18
10
20.2 ± 0.2
40.1 ± 1.0
38.5 ± 2.1
45.6 ± 1.4
47.3 ± 0.8
50
23.3 ± 0.3
67.0 ± 1.8
68.5 ± 0.7
68.9 ± 1.1
69.7 ± 1.6
ResNet-50
10
17.3 ± 1.7
37.3 ± 1.1
29.9 ± 2.2
40.1 ± 0.1
41.3 ± 1.1
50
24.8 ± 0.7
64.3 ± 0.9
67.8 ± 0.3
68.2 ± 1.1
69.0 ± 0.4
ResNet-101
10
17.7 ± 0.9
34.2 ± 1.7
31.3 ± 1.3
36.5 ± 1.4
37.8 ± 1.6
50
21.2 ± 0.2
62.7 ± 1.6
59.1 ± 0.7
63.1 ± 1.3
63.5 ± 1.5
Appendix
Table 7: Performance comparison with soft-label-based baselines on ImageWoof under different IPC and backbone settings. The best results in each setting are highlighted in bold.
Figure 5: Ablation studies on the hyperparameters used in MST construction, including the number of neighbors k in the kNN graph, the maximum number of sampled latent features per class, and the PCA projection dimension.
Figure 6: Comparison between samples generated by Minimax Diffusion (left) and generated by the proposed MGPO method (right) on ImageWoof (Classification).
Figure 7: Comparison between samples generated by Minimax Diffusion (left) and generated by the proposed MGPO method (right) on ImageNette (Classification).
Figure 8: Comparison between samples generated by Minimax Diffusion (left) and generated by the proposed MGPO method (right) on ImageIDC (Classification).
Figure 9: Samples generated by the proposed MGPO (Object Detection).
Figure 10: Samples generated by the proposed MGPO (Segmentation).
Dataset distillation enables efficient training by distilling the information of large-scale datasets into significantly smaller synthetic datasets. Diffusion based paradigms have emerged in recent years, offering novel perspectives for dataset distillation. However, they typically necessitate additional fine-tuning stages, and effective guidance mechanisms remain underexplored. To address these limitations, we rethink diffusion based dataset distillation and propose a Dual Matching Guided Diffusion (DMGD) framework, centered on efficient training-free guidance. We first establish Semantic Matching via conditional likelihood optimization, eliminating the need for auxiliary classifiers. Furthermore, we propose a dynamic guidance mechanism that enhances the diversity of synthetic data while maintaining semantic alignment. Simultaneously, we introduce an optimal transport (OT) based Distribution Matching approach to further align with the target distribution structure. To ensure efficiency, we develop two enhanced strategies for diffusion based framework: Distribution Approximate Matching and Greedy Progressive Matching. These strategies enable effective distribution matching guidance with minimal computational overhead. Experimental results on ImageNet-Woof, ImageNet-Nette, and ImageNet-1K demonstrate that our training-free approach achieves significant improvements, outperforming state-of-the-art (SOTA) methods requiring additional fine-tuning by average accuracy gains of 2.1%, 5.4%, and 2.4%, respectively.
Qichao Wang, Yunhong Lu, Hengyuan Cao +2
Zhejiang University · Shanghai Institute for Advanced Study-Zhejiang University · Shanghai Institute for Mathematics and Interdisciplinary Sciences
Dataset distillation (DD) aims to compress large-scale datasets into compact synthetic sets while preserving training efficacy. However, existing studies mainly focus on image classification, leaving dense prediction tasks such as semantic segmentation largely underexplored. In this work, we identify three key challenges for segmentation DD: (i) long-tailed class imbalance, (ii) the need for strict pixel-wise alignment between images and dense labels, and (iii) the high computational cost of optimizing high-resolution data with complex models. To address these challenges, we propose D3S2, a Diffusion-guided Dataset Distillation framework for Semantic Segmentation. Our method adopts a two-stage design. In Class-Balanced Mask Selection, we construct a representative mask set via a greedy strategy that prioritizes underrepresented classes. In Diffusion-Guided Image Synthesis, we employ a pretrained layout-to-image diffusion model to generate images conditioned on the selected masks, naturally ensuring spatial alignment. To further enhance the training utility of synthesized data, we introduce guided diffusion sampling with two complementary objectives: a segmentation-consistency loss for pixel-level alignment, and a class-wise feature matching loss for aligning per-class feature statistics across layers. Extensive experiments demonstrate the superiority of D3S2. Notably, at an extremely compression rate of 1%, our method achieves 24.99% and 35.49% mIoU on ADE20K and COCO-Stuff with Mask2Former (Swin-S), outperforming random selection by 9.34% and 5.70%, respectively. Our code is available at https://github.com/zwj084/D3S2.
Dataset distillation aims to synthesize compact datasets that can approximate the performance of full-data training while significantly reducing computational and storage costs. However, diffusion-based distillation methods often struggle to preserve structural coherence and generalization, especially in visually complex domains. This issue often stems from latent prototypes that are weakly aligned with class-discriminative regions and contaminated by irrelevant background, thereby degrading generation quality and generalization. To address this limitation, we propose a saliency-driven distillation framework that constructs class-discriminative latent prototypes to enhance representativeness and generalization. The framework proceeds in two stages: (1) ensemble Grad-CAM++ saliency is used to construct prototypes emphasizing class-discriminative regions, and (2) hard-prototype refinement is then applied to construct challenging yet class-consistent prototypes, thereby enhancing discriminability and diversity. Importantly, the diffusion backbones (e.g., LDM and DiT) remain frozen; only lightweight classifiers used for saliency extraction are trained. Extensive experiments across multiple benchmarks demonstrate consistent performance improvements over strong baselines. Code will be released.
Yawen Zou, Wenqi Cai, Guang Li +3
University of Toyama · Hokkaido University · University of Fukui