Synthetic data can scale training supervision when real-world data are limited, but noise and distribution mismatch can reduce its value. Existing synthetic data selection methods often emphasize fidelity or diversity rather than the learner's evolving needs. We propose FROST, an online framework that estimates synthetic-data utility through gradient feedback anchored in real training data. It calibrates batch utility against recent history to determine when filtering is needed and filters samples only in out-of-band batches to determine what to retain, without an external verifier or held-out validation set. Experiments on two public benchmarks for image classification and LLM fine-tuning for text-to-SQL show that FROST filters out around 20--30% of the synthetic data while improving real-task performance compared with training on the full synthetic data pool. We further apply FROST during training in a large-scale industrial ads re-ranking system, achieving significant performance gains over a highly optimized production baseline, demonstrating its effectiveness and generalizability.
Figures & tables
Figure 1: Overview of FROST . The framework \scriptsize1⃝ estimates each synthetic sample’s utility through gradient alignment with a smoothed real-training reference, \scriptsize2⃝ aggregates these scores and calibrates batch utility against recent training history to retain moderate-utility batches directly, and \scriptsize3⃝ applies IQR-based sample filtering to out-of-band batches.
FLUX
SANA
SD1.4
Method
Drop
Acc. ↑
Δ↑
Drop
Acc. ↑
Δ↑
Drop
Acc. ↑
Δ↑
Real only
100
76.04
–
100
76.04
–
100
76.04
–
Real + All Synthetic
0
80.31
0.00
0
79.32
0.00
0
79.46
0.00
Random
25.76
80.33
+0.02
26.01
79.61
+0.29
24.58
79.34
−0.12
DS3
25.76
80.01
−0.30
26.01
79.77
+0.45
24.58
79.44
−0.02
Covariance Matching
25.76
79.92
−0.39
26.01
79.21
−0.11
24.58
79.21
−0.25
Table 1: CIFAR-100 classification with synthetic training data. Δ is the accuracy change relative to Real + All Synthetic for the same generator.
Figure 2: Test accuracy over following 30 steps.
EX by difficulty (%) ↑
Method
Drop
EX ↑
Δ EX ↑
Easy
Medium
Hard
Extra Hard
Real only
100
57.1
–
79.4
59.6
43.7
30.7
Real + All Synthetic
0
64.8
0.0
83.5
67.0
56.9
39.2
Random
20.4
64.1
−0.7
83.9
67.5
51.7
38.6
GradNorm
20.4
64.2
−0.6
82.3
67.7
55.7
36.7
GREATS
20.4
65.3
+0.5
81.9
68.4
57.5
40.4
Table 2: Spider 1.0 fine-tuning with synthetic data pool, evaluated on the dev split. EX is execution accuracy (%), reported overall and by difficulty. Δ EX is the overall EX change relative to Real + All Synthetic. All methods retain the real training data; selection methods use a matched drop ratio.
Figure 3: Two-stage ablation on Spider 1.0 dev, measured against the real-only model.
Training data / selection
Drop (%)
Relative NE change (%) ↓
Real data only (reference)
100
–
+ All synthetic data
0
+0.211↑
+ Random
30
+0.118↑
+ Random
70
+0.023↑
FROST
36
−0.096↓
Table 3: Ads re-ranking results relative to real-only training. Drop is the synthetic-data drop ratio; all real data are retained. Red ↑ denotes higher NE (worse); green ↓ denotes lower NE (better).
Figure 4: Sensitivity to the batch utility band on the ads model. Relative NE change (%) uses the real-only reference; lower is better. Dashed lines mark real-only training ( 0% ) and all synthetic data ( +0.211% ). (a) Narrowing the symmetric z -band does not consistently improve NE. (b) The asymmetric band achieves lower NE at a similar drop ratio.
Figure 5: Two-stage ablation. Percentage denotes the synthetic data drop ratio.
Appendix figures & tables3 assets
Supplementary material from the paper’s appendix.
Appendix
Figure 6: Examples from the image-generation pipeline. Rows show apple, bicycle, and castle. The first column contains real CIFAR-100 images; the remaining columns are SANA-1.6B outputs conditioned on the cleaned LLaVA caption (v0) and three Qwen rewrites (v1–v3). Generated images are downsampled to 32×32 with Lanczos and enlarged with nearest-neighbor interpolation for display; real images are enlarged from their native 32×32 resolution. The bicycle captions illustrate how rewriting changes the scene while preserving the subject.
Computation
MACs per step
Regular forward
(∣Rt∣+m)(Fbackbone+Cd)
Regular backward
≈2(∣Rt∣+m)(Fbackbone+Cd)
Real-reference aggregation
∣Rt∣Cd
Synthetic utility scoring
mCd
EMA update
O(Cd)
Appendix
Table 4: Leading arithmetic costs for regular ResNet-18 training and head-level utility computation. The latter assumes that features and logit gradients are already available.
EX by difficulty (%) ↑
Variant
Drop
EX ↑
Δ EX ↑
Easy
Medium
Hard
Extra Hard
Full FROST
20.4
65.7
0.0
83.5
68.6
54.6
42.8
Batch gate only
80.0
60.7
−5.0
79.0
63.5
53.4
33.7
Sample IQR only
32.2
63.1
−2.6
80.6
65.9
56.3
36.1
Appendix
Table 5: Two-stage ablation on Spider 1.0 dev. Full FROST is the reference. Drop is the synthetic-data drop ratio (%); EX is execution accuracy (%). Δ EX is the overall EX change relative to full FROST , in percentage points (pp). All variants retain the real training data.
Synthetic data are increasingly used to scale LLM training, yet more synthetic data do not necessarily produce better models. Useful synthetic data must add information relevant to the target task without introducing errors that offset their benefit, and the value of an example can change as the training set grows. We develop a linear theory that characterizes this tradeoff and determines where synthetic data are useful, how much should be added, and the marginal value of adding one example to an existing set. The analysis shows the conditions when input coverage alone is sufficient and when synthetic errors must also be considered. Guided by these results, we introduce \emph{Training-Aware Target Coverage} (TATC), a synthetic data selection method for LLM fine-tuning. TATC identifies candidates whose training effects are beneficial to the target task and selects among them to expand coverage of target-relevant directions not already represented by the available data. Experiments on text and image data verify the linear theory. With mathematical reasoning tasks, TATC selects synthetic solutions for fine-tuning Qwen2.5-Math-1.5B-Instruct and outperforms alternative synthetic-data selection methods on GSM8K across selection budgets. In summary, we provide a principled approach to synthetic data selection by quantifying and maximizing its value to the target task.
Yang Ba, Michelle V. Mancenido, Rong Pan
School of Computing and Augmented Intelligence, Arizona State University · School of Mathematical and Natural Sciences, Arizona State University
Synthetic data is useful only when the added samples fill missing parts of the training distribution that matter for the downstream task. We introduce LiBaGS, a lightweight, generator-agnostic method for targeted synthetic training data selection. LiBaGS scores candidate synthetic samples by combining decision-boundary proximity, predictive uncertainty, real-data density, and support validity, so that selected samples are both informative and likely to remain on the real data manifold. We then use a boundary-gap allocation rule that targets sparse but realistic decision-boundary neighborhoods, rather than simply adding more data or selecting only the most uncertain candidates. LiBaGS also learns when enough synthetic samples have been added through a marginal-value stopping rule, assigns softer labels near ambiguous boundaries, and uses a diversity objective to avoid redundant near-duplicate selections. Experiments show that LiBaGS improves accuracy over classical oversampling, hard augmentation, uncertainty and density ablations, and targeted-generation selection criteria.
Abhishek Moturu, Anna Goldenberg, Babak Taati
Department of Computer Science University of Toronto The Hospital for Sick Children UHN KITE Research Institute T-CAIREM Vector Institute · Department of Computer Science Department of Laboratory Medicine and Pathobiology University of Toronto The Hospital for Sick Children T-CAIREM Vector Institute · Department of Computer Science Institute of Biomedical Engineering University of Toronto Rehabilitation Sciences Institute UHN KITE Research Institute Vector Institute
While synthetic data generation with large language models (LLMs) is widely used in post-training pipelines, existing approaches typically generate full outputs before applying quality filters, leading to substantial token waste on samples that are ultimately discarded. To address this, we propose Multi-Stage In-Flight Rejection (MSIFR), a lightweight, training-free framework that detects and terminates low-quality generation trajectories at intermediate checkpoints before they reach full completion. MSIFR decomposes the generation process into sequential stages and applies fast rule-based validators to identify arithmetic inconsistencies, hallucination patterns, and formatting violations, enabling early rejection of faulty samples. We formalize in-flight rejection as a sequential decision process and show that any non-trivial discard policy reduces expected token consumption, with stage-wise savings increasing when rejection occurs earlier in the generation pipeline. We further demonstrate that conditional utility estimates form a martingale, ensuring that early, in-flight rejection does not bias the expected utility of retained samples. Across five instruction-tuned models and seven reasoning benchmarks, MSIFR reduces token consumption by 11%-77% as a standalone method, and up to 78.2% when combined with early-exit methods, while preserving or improving evaluation accuracy. These results confirm that MSIFR provides a practical mechanism for improving the efficiency of LLM-based synthetic data generation without additional training or architectural changes.
Anjir Ahmed Chowdhury, Syed Zawad, Feng Yan
Department of Computer Science University of Houston · IBM Research