Selecting synthetic degradations for dense prediction requires an estimate of their training utility, the generalization gain they bring under a finite training budget. Clean and degraded twins share content and labels, suggesting a score based on how much short training reduces the excess error caused by degradation. However, this gap can also shrink when clean performance deteriorates. Measuring the improvement on degraded images alone avoids that confound, but it still credits progress that the same amount of clean training would have produced. We propose the \textbf{controlled Reducible Degradation Gap} (cRDG) for regions defined by degradation type and severity. From a common checkpoint, cRDG runs two budget-matched probes that differ only in one augmentation slot, which holds either a synthetic degradation or a clean augmentation. The score is the gain on held-out degraded images relative to the clean-control probe. Clean harm is a separate feasibility constraint. cRDG reveals a correctable severity band in which training on the degradation yields high controlled gain under the available budget, and the band moves with the predictor, the starting checkpoint, and the training budget. \textbf{Curation of Reducible Bands} (\method) uses cRDG to select synthetic data without changing the predictor. On semantic segmentation and salient object detection, \method{} improves representative predictors under matched synthetic-data budgets and training schedules, extends to existing data-generation pipelines, and preserves clean performance. Code and supporting materials will be publicly released.
Figures & tables
Figure 1 : Why Matched Clean Control Matters. (a) Paired gap reduction can be inflated by clean forgetting. (b) Held-out degraded loss drop avoids this inflation but still credits generic training progress. (c) Our matched-control formulation underlying cRDG measures degraded gain Gg beyond equally long clean training, while checking clean harm Hg separately.
Figure 2 : Overview of CuRB . (a) Degradation-severity regions are generated and validity-filtered. (b) Two matched probes start from θ0 and differ only in the augmentation slot, where ⊕ denotes the union of the clean-replay stream and the slot stream. (c) Held-out evaluation yields degraded gain Gg and clean harm Hg . Averaging Gg over the two split directions gives cRDG . Hg determines feasibility, while the correctable band is diagnostic only. (d) Feasible positive-score regions are ranked and fill the budget Bsel in rank order, with k -center sampling for the last partial region. (e) Final training restarts from θ0 on clean data, with S⋆ in the synthetic slots under the fixed schedule.
Table 1 : Evaluation on SOD Benchmarks following the official protocol ( Section 4.1 ). At the same augmentation budget, comparing uncurated sampling with CuRB isolates the gain from curation.
ACDC mIoU ↑
Dark Zurich mIoU ↑
Cityscapes
Data-generation host
Published host setting
Original
+ CuRB
Original
+ CuRB
val Δ↑
single-generator hosts
ISSA [ 17 ]
SegFormer
52.45
53.92
27.39
28.76
+0.24
H-Weather [ 13 ]
SegFormer
49.21
50.63
23.44
24.65
+0.36
Gen4Seg [ 31 ]
SegFormer
51.04
52.39
25.63
27.28
+0.45
composite DG host
Table 2 : CuRB as a plug-in to existing data-generation settings. Published references for ISSA, H-Weather, and RobustNet+ISSA are taken from Li et al. [17] . Gen4Seg follows Yin et al. [31] (its Dark Zurich value is obtained with the released pipeline). Our “+ CuRB ” runs apply CuRB under the corresponding host setting and synthetic-data budget. The composite host stacks RobustNet [ 6 ] with ISSA on DeepLabV3+. Host-specific region construction: Appendix D .
ACDC in-family
in-fam.
Snow
Dark
Cityscapes
Method
Fog
Night
Rain
mean
(held out)
Zurich
val
Source-only
60.61
28.42
50.35
46.46
48.79
24.08
67.84
matched curation protocol: same pool, gate, Bsel , objective, and schedule
All-gen. (uniform, Bsel )
63.82
31.51
52.78
49.37
50.15
26.55
68.21
Realism (NR-IQA / CLIPScore)
63.34
30.97
53.02
49.11
50.31
26.12
68.44
Highest loss / uncertainty
64.05
31.88
52.41
49.45
49.97
26.71
67.93
Table 3 : Controlled curation comparison on Cityscapes → ACDC / Dark Zurich. All methods share the candidate pool, validity gate, Bsel , final-training objective, and schedule, differing in their sample-selection policies. Fog, night, and rain are treated as the in-family conditions. Snow is held out as an unseen family, with no snow samples entering the candidate pool, curation, or final training.
Figure 3 : Qualitative comparison. (a) Degraded SOD with NUN. (b) adverse-condition SemSeg with ISSA SegFormer. Baseline and + CuRB predictions are shown for both tasks.
Table 7
Figure 4 : Training utility and the correctable band. (a) Per-bin retraining gains (top) and scores (bottom) for low light. Shading marks the predicted band. (b) Empirical severity centroids at three final budgets with scaled- K cRDG predictions overlaid (paired-bootstrap 95% CI). (c) cRDG centroids across probe budgets, checkpoints, and predictors. Error bars are s.d. over three partitions.
Adapted valuation baselines
Twin-based scores
Metric
Highest loss / uncertainty
RHO-Loss style [ 19 ]
Grad. contrib. [ 34 ]
Δ0 (gap)
RDG (uncontrolled)
Dg (held-out drop)
CuRB ( cRDG )
ρglob with Ureal↑
0.25
0.49
0.52
0.39
0.61
0.70
0.79
ρmacro with Ureal↑
0.17
0.35
0.37
0.27
0.43
0.56
0.74
ρDZ with UDZ↑
0.21
0.43
0.45
0.33
0.53
0.62
0.72
in-fam. mIoU ↑
49.45
49.80
50.00
49.28
50.39
51.02
51.97
Table 6 : Region-level score comparison under matched pool, gate, diversity, budget, objective, and schedule (Cityscapes → ACDC, ISSA SegFormer). ρglob is Spearman correlation with Ureal over all 32 regions, ρmacro averages within-family correlations, and ρDZ evaluates the same ranking against per-region retraining gains on Dark Zurich. Δ0 is the unmastered gap and Dg the held-out loss drop.
Appendix figures & tables3 assets
Supplementary material from the paper’s appendix.
Appendix
Degradation
gate prec.
gate recall
accept. s1 – s4
accept. s5 – s8
Low light
0.94
0.91
0.99
0.95
Backlight
0.92
0.89
0.99
0.92
Fog
0.95
0.92
1.00
0.96
Rain
0.91
0.88
0.98
0.90
Appendix
Tab. S1 : Validity-gate audit on 500 sampled candidates per family. Structural drift is the positive class for gate precision and recall. Acceptance remains high at severe levels ( 0.90 – 0.96 ), indicating that the gate does not simply remove stronger degradations.
ϵ
0
0.5SE
1SE (default)
2SE
feasible-region fraction
0.41
0.63
0.78
0.91
in-family mIoU ↑
51.43
51.86
51.97
51.79
Cityscapes mIoU ↑
70.21
70.24
70.13
69.82
Appendix
Tab. S2 : Sensitivity to the clean-retention tolerance ϵ . Degraded-domain performance is stable around the default, while loosening the constraint permits a small clean-performance drop.
Quantity
Value
Regions M (four families × eight severity bins)
32
Split directions R
2
Probe horizon K
1/80 of the final schedule
Treatment probes ( MR )
64
Shared control probes ( R )
2
Probe-step cost / final-step cost (measured)
≈0.71
Appendix
Tab. S3 : Curation cost. Operational selection uses one A∣B partition. The region-independent control adds only two probes, and the normalized probe-update cost is (66/80)×0.71≈0.59 of one final-training run, under the K=T/80 scaling rule.
Synthetic images are increasingly used to augment scarce real data for object detection. However, not all synthetic sets help equally, and the only way to know a set's value is to train a detector on it, which is slow and demands dense annotation. We ask whether a training-free metric can instead rank candidate synthetic training sets by their downstream utility. Existing image-set metrics such as FID, KID, and MMD compare two feature distributions with a single global statistic, which we show is mis-specified for detection-data selection in two ways: it is blind to per-image composition (object count, box scale, class mix), and even at fixed composition its global averaging washes out the appearance differences that separate high-mAP pools from low-mAP ones. We propose Conditional-Composition Domain Match (CCDM), which converts any feature-space distance into a composition-stratified comparison, matching candidate and target within metadata-defined strata without training a detector. On COCO and VisDrone-DET, the best CCDM variant ranks 19 candidate training sets in strong agreement with YOLOv8 mAP (Spearman \r{ho} = 0.97 and 0.96), outperforming FID, KID, and MMD. Furthermore, CCDM holds when reference metadata comes from detector pseudo-labels rather than ground-truth boxes.
Recent generative models can produce high-quality synthetic images, offering scalable training training data for data-hungry models. Existing approaches to exploiting this potential typically involve 1) training or fine-tuning generators, or 2) using lightweight post-hoc adaptation like prompt engineering or inference-time guidance, making them generator-specific and expertise-intensive. We study a complementary question: given a fixed pool of generated images, can downstream utility be improved purely by selecting an informative subset? The answer is yes. We show that effective selection must counter a structural bias of modern generators: they tend to over-produce canonical modes of each class while underrepresenting intra-class variation. Building on this insight, we split each real class into a canonical Homogeneous (HO) subset and a non-redundant Heterogeneous (HE) subset, then score synthetic images by a fidelity-diversity criterion that rewards semantic alignment while penalizing canonical redundancy. The method is generator-agnostic and requires no retraining. Across multiple benchmarks, it consistently outperforms state-of-the-art data selection baselines and matches the real-data performance with up to 40% fewer synthetic samples. The same criterion remains effective when applied on top of stronger task-tuned generators, with gains on both classification and segmentation tasks. Post-generation selection is therefore not a substitute for better generators, but a complementary mechanism for improving the utility of synthetic data.
Disheng Liu, Tuo Liang, Chaoda Song +1
Department of Computer and Data Sciences Case Western Reserve University Cleveland, OH, USA
Synthetic data now is essential to LLM training, used to strengthen advanced capabilities such as autonomous and long-horizon task execution. Yet recent work shows that training on it at scale can degrade model generation, making it important to decide what synthetic data is worth training on. While current data curation practices do so with individual-level signals (i.e., estimates of each data sample's training utility in isolation), across pre-training and post-training settings we show that this is insufficient for synthetic data, and that group-level signals (i.e., estimates of utility that account for interactions among data samples) are necessary for effective data curation. First, we show that individual-level signals are blind to how samples jointly affect training: synthetic datasets with different compositions can be indistinguishable under individual-level influence yet differ sharply under group-level influence, and curating by the latter yields better downstream performance, particularly in generative capability. Second, we find that group-level signals matter more as training pipelines become increasingly synthetic: among widely used data curation methods, only those incorporating them improve over baseline, with gains increasing when weights capturing relations among samples are amplified. Finally, we translate these findings into practice -- for model developers under a compute budget, we offer a cheap diagnostic that prioritizes which groups of synthetic data most need group-level estimation, recovering much of the benefit of full group-level scoring at a fraction of the compute cost.
Cathy Jiao, Chenyan Xiong
Language Technologies Institute, Carnegie Mellon University