cs.CVOct 4, 2026

MGPO: Manifold-Guided Diffusion Alignment for Task-Aware Dataset Distillation

Authors: Yunyi Chen, Chenru Wang, Xinyi Ye, Zexin Zheng, Chi Zhang

Organizations: AGI Lab, Westlake University · Eindhoven University of Technology · Guangdong University of Finance and Economics

Abstract

Diffusion-based dataset distillation (DD) suffers from a fundamental objective mismatch: likelihood-driven diffusion models prioritize density approximation over the discriminative decision boundaries required for downstream tasks. Beyond semantic mismatch, relying solely on density also leads to geometric coverage loss, where generated samples collapse into a few high-density modes and fail to cover the manifold's structural diversity. We propose Manifold-Guided Policy Optimization (MGPO), which reformulates DD as a multi-objective reinforcement learning problem and achieves Dual-Space Alignment via a pixel-space discriminative reward and a latent-space geometric reward guided by a class-wise Minimum Spanning Tree (MST). The discriminative reward enforces class separability, while the MST-based geometric reward encourages generated latents to cover a sparse geometric skeleton of each class, jointly addressing both failure modes. We further provide an idealized analysis that motivates the MST-based reward, including a Hausdorff approximation bound and a subsampling bound independent of the dataset size. The reward-modular design extends to structured tasks such as object detection and segmentation by substituting the frozen task reward model. Extensive experiments show MGPO consistently outperforms existing methods, including a +8.0% mIoU gain on segmentation under low-budget settings.

Figures & tables

Appendix figures & tables7 assets

Supplementary material from the paper’s appendix.

Appendix

Explore similar work

CardsList
  1. DMGD: Train-Free Dataset Distillation with Semantic-Distribution Matching in Diffusion Models

    May 5, 2026Qichao Wang, Yunhong Lu, Hengyuan Cao +2Diffusion-Based Dataset DistillationDataset Distillation

  2. Dataset Distillation Based on Saliency-Driven Prototype Alignment

    Jul 28, 2026Yawen Zou, Wenqi Cai, Guang Li +3Diffusion-Based Dataset DistillationDataset Distillation