Semi-Supervised Conditional Diffusion via Label Augmentation
Organizations: aDepartment of Applied Mathematics, The Hong Kong Polytechnic University, Hong Kong, China · bSchool of Statistics and Data Science, Nankai University, Tianjin, China · cSchool of Statistics, East China Normal University, Shanghai, China · dDepartment of Data Science and AI, The Hong Kong Polytechnic University, Hong Kong, China
Abstract
Conditional diffusion models have become a powerful and flexible framework for learning complex conditional distributions from labeled data. In practice, however, acquiring high-quality labels is costly and time-consuming, leaving large volumes of unlabeled data unused. To address this, we introduce label-augmented conditional diffusion (LACD), a simple and effective approach that incorporates unlabeled examples by assigning them a designated trivial label and performing joint denoising score matching over the augmented dataset. We provide sufficient conditions guaranteeing population-level identifiability of the target conditional distribution under this scheme. Moreover, we establish rigorous statistical guarantees: when sufficiently many unlabeled samples are available, the sampling distribution produced by LACD converges strictly faster than the purely supervised estimator in total variation distance, and at least as fast in Wasserstein-1 distance. Extensive experiments on synthetic, image, and tabular benchmarks corroborate our theory and show substantial gains in sample efficiency and generative performance compared with the purely supervised estimator.