cs.CVSep 28, 2026

PSM: Dataset Distillation Based on Precise Statistical Matching by Difficulty

Authors: Hongxu Ma, Guang Li, Shijie Wang, Dongzhan Zhou, Suorong Yang, Baoli Sun, Takahiro Ogawa, Miki Haseyama, +1 more

Organizations: Zhejiang University · Hokkaido University · The University of Queensland · Shanghai Artificial Intelligence Laboratory · National University of Singapore · Dalian University of Technology

Abstract

Dataset distillation (DD) condenses a large original dataset into a small distilled dataset with high training utility. Decoupled statistical matching methods substantially reduce distillation time and memory overhead while achieving strong performance. However, they typically supervise all distilled samples using running statistics estimated from the entire original dataset. These statistics mainly capture the average feature distribution while overlooking differences in sample difficulty, limiting their ability to characterize the difficulty structure of the original data. To address this issue, we propose Precise Statistical Matching (PSM) by difficulty. After pretraining, PSM uses the Global Precision Score (GPS) to estimate image difficulty, ranks the samples within each class, and partitions each class into IPC (images per class) difficulty groups. During distillation, Statistics Updated Again (SUA) updates the teacher's batch normalization (BN) running statistics through forward passes on original samples from each group, providing difficulty-specific supervision for the corresponding distilled batch. Meanwhile, Initial Sample Screening (ISS) initializes distilled samples using original images from the corresponding difficulty group, providing an effective starting point for precise matching. Experiments across multiple datasets and model architectures demonstrate that PSM broadens the difficulty range of distilled samples and improves downstream performance in most evaluated settings. Code will be released.

Figures & tables

Appendix figures & tables12 assets

Supplementary material from the paper’s appendix.

Appendix

Explore similar work

CardsList
  1. Efficient Dataset Distillation for Pre-Trained Self-Supervised Models via Statistical Flow Matching

    Feb 5, 2026Qianxin Xia, Jiawei Du, Yuhan Zhang +6Diffusion-Based Dataset DistillationSelf-Supervised Learning

  2. FD2^2: A Dedicated Framework for Fine-Grained Dataset Distillation

    Mar 26, 2026Hongxu Ma, Guang Li, Shijie Wang +5Diffusion-Based Dataset DistillationDataset Distillation

  3. Dataset Distillation by Influence Matching

    Jul 18, 2026Haoru Tan, Wang Wang, Sitong Wu +5Diffusion-Based Dataset DistillationDataset Distillation