cs.CVOct 7, 2026

Hard, Yet Reducible: Controlled Forward Transfer for Synthetic Degradation Curation

Authors: Chunming He, Kailai Zhou, Jiaming Zuo, Hanqi Liu, Fengyang Xiao, Youwei Pang, Xiaofeng Liu, Weisi Lin, +1 more

Organizations: Tsinghua University, China · X3000 Inspection Co., Ltd · AI4X Team, Nanyang Technological University, Singapore · Yale University, USA

Abstract

Selecting synthetic degradations for dense prediction requires an estimate of their training utility, the generalization gain they bring under a finite training budget. Clean and degraded twins share content and labels, suggesting a score based on how much short training reduces the excess error caused by degradation. However, this gap can also shrink when clean performance deteriorates. Measuring the improvement on degraded images alone avoids that confound, but it still credits progress that the same amount of clean training would have produced. We propose the \textbf{controlled Reducible Degradation Gap} (cRDG) for regions defined by degradation type and severity. From a common checkpoint, cRDG runs two budget-matched probes that differ only in one augmentation slot, which holds either a synthetic degradation or a clean augmentation. The score is the gain on held-out degraded images relative to the clean-control probe. Clean harm is a separate feasibility constraint. cRDG reveals a correctable severity band in which training on the degradation yields high controlled gain under the available budget, and the band moves with the predictor, the starting checkpoint, and the training budget. \textbf{Curation of Reducible Bands} (\method) uses cRDG to select synthetic data without changing the predictor. On semantic segmentation and salient object detection, \method{} improves representative predictors under matched synthetic-data budgets and training schedules, extends to existing data-generation pipelines, and preserves clean performance. Code and supporting materials will be publicly released.

Figures & tables

Appendix figures & tables3 assets

Supplementary material from the paper’s appendix.

Appendix

Explore similar work

Jun 18, 2026cs.CV

Training-Free Metrics for Synthetic Object Detection Data: A Proxy for Detector Performance

Synthetic images are increasingly used to augment scarce real data for object detection. However, not all synthetic sets help equally, and the only way to know a set's value is to train a detector on it, which is slow and demands dense annotation. We ask whether a training-free metric can instead rank candidate synthetic training sets by their downstream utility. Existing image-set metrics such as FID, KID, and MMD compare two feature distributions with a single global statistic, which we show is mis-specified for detection-data selection in two ways: it is blind to per-image composition (object count, box scale, class mix), and even at fixed composition its global averaging washes out the appearance differences that separate high-mAP pools from low-mAP ones. We propose Conditional-Composition Domain Match (CCDM), which converts any feature-space distance into a composition-stratified comparison, matching candidate and target within metadata-defined strata without training a detector. On COCO and VisDrone-DET, the best CCDM variant ranks 19 candidate training sets in strong agreement with YOLOv8 mAP (Spearman \r{ho} = 0.97 and 0.96), outperforming FID, KID, and MMD. Furthermore, CCDM holds when reference metadata comes from detector pseudo-labels rather than ground-truth boxes.
Jul 2, 2026cs.LG

Post-Generation Curation of Synthetic Images via Homogeneous-Heterogeneous Splitting

Recent generative models can produce high-quality synthetic images, offering scalable training training data for data-hungry models. Existing approaches to exploiting this potential typically involve 1) training or fine-tuning generators, or 2) using lightweight post-hoc adaptation like prompt engineering or inference-time guidance, making them generator-specific and expertise-intensive. We study a complementary question: given a fixed pool of generated images, can downstream utility be improved purely by selecting an informative subset? The answer is yes. We show that effective selection must counter a structural bias of modern generators: they tend to over-produce canonical modes of each class while underrepresenting intra-class variation. Building on this insight, we split each real class into a canonical Homogeneous (HO) subset and a non-redundant Heterogeneous (HE) subset, then score synthetic images by a fidelity-diversity criterion that rewards semantic alignment while penalizing canonical redundancy. The method is generator-agnostic and requires no retraining. Across multiple benchmarks, it consistently outperforms state-of-the-art data selection baselines and matches the real-data performance with up to 40% fewer synthetic samples. The same criterion remains effective when applied on top of stronger task-tuned generators, with gains on both classification and segmentation tasks. Post-generation selection is therefore not a substitute for better generators, but a complementary mechanism for improving the utility of synthetic data.
Sep 30, 2026cs.CL

Effective Synthetic Data Curation Requires Group-Level Signals

Synthetic data now is essential to LLM training, used to strengthen advanced capabilities such as autonomous and long-horizon task execution. Yet recent work shows that training on it at scale can degrade model generation, making it important to decide what synthetic data is worth training on. While current data curation practices do so with individual-level signals (i.e., estimates of each data sample's training utility in isolation), across pre-training and post-training settings we show that this is insufficient for synthetic data, and that group-level signals (i.e., estimates of utility that account for interactions among data samples) are necessary for effective data curation. First, we show that individual-level signals are blind to how samples jointly affect training: synthetic datasets with different compositions can be indistinguishable under individual-level influence yet differ sharply under group-level influence, and curating by the latter yields better downstream performance, particularly in generative capability. Second, we find that group-level signals matter more as training pipelines become increasingly synthetic: among widely used data curation methods, only those incorporating them improve over baseline, with gains increasing when weights capturing relations among samples are amplified. Finally, we translate these findings into practice -- for model developers under a compute budget, we offer a cheap diagnostic that prioritizes which groups of synthetic data most need group-level estimation, recovering much of the benefit of full group-level scoring at a fraction of the compute cost.