Synthetic data can scale training supervision when real-world data are limited, but noise and distribution mismatch can reduce its value. Existing synthetic data selection methods often emphasize fidelity or diversity rather than the learner's evolving needs. We propose FROST, an online framework that estimates synthetic-data utility through gradient feedback anchored in real training data. It calibrates batch utility against recent history to determine when filtering is needed and filters samples only in out-of-band batches to determine what to retain, without an external verifier or held-out validation set. Experiments on two public benchmarks for image classification and LLM fine-tuning for text-to-SQL show that FROST filters out around 20--30% of the synthetic data while improving real-task performance compared with training on the full synthetic data pool. We further apply FROST during training in a large-scale industrial ads re-ranking system, achieving significant performance gains over a highly optimized production baseline, demonstrating its effectiveness and generalizability.
Figures & tables
Figure 1: Overview of FROST . The framework \scriptsize1⃝ estimates each synthetic sample’s utility through gradient alignment with a smoothed real-training reference, \scriptsize2⃝ aggregates these scores and calibrates batch utility against recent training history to retain moderate-utility batches directly, and \scriptsize3⃝ applies IQR-based sample filtering to out-of-band batches.
FLUX
SANA
SD1.4
Method
Drop
Acc. ↑
Δ↑
Drop
Acc. ↑
Δ↑
Drop
Acc. ↑
Δ↑
Real only
100
76.04
–
100
76.04
–
100
76.04
–
Real + All Synthetic
0
80.31
0.00
0
79.32
0.00
0
79.46
0.00
Random
25.76
80.33
+0.02
26.01
79.61
+0.29
24.58
79.34
−0.12
DS3
25.76
80.01
−0.30
26.01
79.77
+0.45
24.58
79.44
−0.02
Covariance Matching
25.76
79.92
−0.39
26.01
79.21
−0.11
24.58
79.21
−0.25
Table 1: CIFAR-100 classification with synthetic training data. Δ is the accuracy change relative to Real + All Synthetic for the same generator.
Figure 2: Test accuracy over following 30 steps.
EX by difficulty (%) ↑
Method
Drop
EX ↑
Δ EX ↑
Easy
Medium
Hard
Extra Hard
Real only
100
57.1
–
79.4
59.6
43.7
30.7
Real + All Synthetic
0
64.8
0.0
83.5
67.0
56.9
39.2
Random
20.4
64.1
−0.7
83.9
67.5
51.7
38.6
GradNorm
20.4
64.2
−0.6
82.3
67.7
55.7
36.7
GREATS
20.4
65.3
+0.5
81.9
68.4
57.5
40.4
Table 2: Spider 1.0 fine-tuning with synthetic data pool, evaluated on the dev split. EX is execution accuracy (%), reported overall and by difficulty. Δ EX is the overall EX change relative to Real + All Synthetic. All methods retain the real training data; selection methods use a matched drop ratio.
Figure 3: Two-stage ablation on Spider 1.0 dev, measured against the real-only model.
Training data / selection
Drop (%)
Relative NE change (%) ↓
Real data only (reference)
100
–
+ All synthetic data
0
+0.211↑
+ Random
30
+0.118↑
+ Random
70
+0.023↑
FROST
36
−0.096↓
Table 3: Ads re-ranking results relative to real-only training. Drop is the synthetic-data drop ratio; all real data are retained. Red ↑ denotes higher NE (worse); green ↓ denotes lower NE (better).
Figure 4: Sensitivity to the batch utility band on the ads model. Relative NE change (%) uses the real-only reference; lower is better. Dashed lines mark real-only training ( 0% ) and all synthetic data ( +0.211% ). (a) Narrowing the symmetric z -band does not consistently improve NE. (b) The asymmetric band achieves lower NE at a similar drop ratio.
Figure 5: Two-stage ablation. Percentage denotes the synthetic data drop ratio.
Appendix figures & tables3 assets
Supplementary material from the paper’s appendix.
Appendix
Figure 6: Examples from the image-generation pipeline. Rows show apple, bicycle, and castle. The first column contains real CIFAR-100 images; the remaining columns are SANA-1.6B outputs conditioned on the cleaned LLaVA caption (v0) and three Qwen rewrites (v1–v3). Generated images are downsampled to 32×32 with Lanczos and enlarged with nearest-neighbor interpolation for display; real images are enlarged from their native 32×32 resolution. The bicycle captions illustrate how rewriting changes the scene while preserving the subject.
Computation
MACs per step
Regular forward
(∣Rt∣+m)(Fbackbone+Cd)
Regular backward
≈2(∣Rt∣+m)(Fbackbone+Cd)
Real-reference aggregation
∣Rt∣Cd
Synthetic utility scoring
mCd
EMA update
O(Cd)
Appendix
Table 4: Leading arithmetic costs for regular ResNet-18 training and head-level utility computation. The latter assumes that features and logit gradients are already available.
EX by difficulty (%) ↑
Variant
Drop
EX ↑
Δ EX ↑
Easy
Medium
Hard
Extra Hard
Full FROST
20.4
65.7
0.0
83.5
68.6
54.6
42.8
Batch gate only
80.0
60.7
−5.0
79.0
63.5
53.4
33.7
Sample IQR only
32.2
63.1
−2.6
80.6
65.9
56.3
36.1
Appendix
Table 5: Two-stage ablation on Spider 1.0 dev. Full FROST is the reference. Drop is the synthetic-data drop ratio (%); EX is execution accuracy (%). Δ EX is the overall EX change relative to full FROST , in percentage points (pp). All variants retain the real training data.
Department of Computer Science University of Toronto The Hospital for Sick Children UHN KITE Research Institute T-CAIREM Vector Institute · Department of Computer Science Department of Laboratory Medicine and Pathobiology University of Toronto The Hospital for Sick Children T-CAIREM Vector Institute · Department of Computer Science Institute of Biomedical Engineering University of Toronto Rehabilitation Sciences Institute UHN KITE Research Institute Vector Institute