Training-Aware Target Coverage for Synthetic Data Selection
Authors: Yang Ba, Michelle V. Mancenido, Rong Pan
Organizations: School of Computing and Augmented Intelligence, Arizona State University · School of Mathematical and Natural Sciences, Arizona State University
Synthetic data are increasingly used to scale LLM training, yet more synthetic data do not necessarily produce better models. Useful synthetic data must add information relevant to the target task without introducing errors that offset their benefit, and the value of an example can change as the training set grows. We develop a linear theory that characterizes this tradeoff and determines where synthetic data are useful, how much should be added, and the marginal value of adding one example to an existing set. The analysis shows the conditions when input coverage alone is sufficient and when synthetic errors must also be considered. Guided by these results, we introduce \emph{Training-Aware Target Coverage} (TATC), a synthetic data selection method for LLM fine-tuning. TATC identifies candidates whose training effects are beneficial to the target task and selects among them to expand coverage of target-relevant directions not already represented by the available data. Experiments on text and image data verify the linear theory. With mathematical reasoning tasks, TATC selects synthetic solutions for fine-tuning Qwen2.5-Math-1.5B-Instruct and outperforms alternative synthetic-data selection methods on GSM8K across selection budgets. In summary, we provide a principled approach to synthetic data selection by quantifying and maximizing its value to the target task.
Figures & tables
Figure 1: Risk and value of synthetic data. The theory in one view. (a) A synthetic set is valuable when target-relevant input information outweighs random label noise and systematic label error. (b) This tradeoff determines whether the optimal amount along a direction is zero, finite, or unbounded. (c) Candidate value depends on both input coverage relative to the current set and label error.
Figure 2: Candidate value and synthetic-data amount with real inputs and controlled targets. (a) Candidate-level test on SetFit/subj. Coverage alone does not distinguish beneficial from harmful candidates, whereas the complete marginal closely predicts the risk reduction from explicit refitting. (b) Amount selection with ResNet-50 features on CIFAR-10. The amount predicted from a labeled real probe closely tracks the test-optimal amount and decreases as pseudo-label reliability falls, while the input-only rule selects the full candidate pool. Gray segments indicate amounts within 1% of the test minimum.
Figure 3: (a) Exact match at B=128 on the GSM8K holdout. The dashed line marks the checkpoint trained on real data. (b) Exact match across selection budgets. Each method selects nested subsets, with each subset trained independently from the same checkpoint. Error bars show one standard deviation over three independent training runs.
Appendix figures & tables4 assets
Supplementary material from the paper’s appendix.
Appendix
Partition
Examples
Role
Drep
1,500
Fit PCA and the reference linear target
Dcal
500
Estimate the residual variance scale
Dr
1,200
Define the real training information
Dq
1,000
Estimate the target moment Q
Ps
3,800
Supply candidate inputs
Appendix
Table 1: Data partitions for the text experiment.
Score
Metric
Value
Coverage
AUROC
0.519
Complete marginal
AUROC
0.99
Complete marginal
Correlation
0.99
Complete marginal
Sign agreement
99.2%
Appendix
Table 2: Agreement with observed risk reductions in the text experiment.
Label model
Label accuracy
Predicted ratio
Test optimum ratio
12,000 training examples
0.889
7.0
8.0
400 training examples
0.836
0.90
0.95
12 coordinates removed
0.604
0
0
Appendix
Table 3: Mean label accuracy and ratios of synthetic to real data for the three conditions in Figure 2 (b).
Linear quantity or role
TATC quantity
Status
Correct interpretation
Real information Ar
A0
Approximation
Summarizes prompt directions already represented by real data; ρI is numerical stabilization.
Target moment Q
Q
Sample estimate
Estimates target relevance only in the checkpoint prompt representation.
Selected information AS
∑i∈Sϕiϕi⊤
Exact in representation
Each selected input adds a rank-one prompt-feature contribution.
Clean coverage marginal qj/cj
aj(S)
Exact in representation
Exact decrease in the downstream-weighted inverse-trace criterion.
Systematic label error (dS,u)
Response-dependent gradient and AdamW update
Approximation
For squared loss, response error changes the gradient by −(x⊤δ)x ; in an LLM the full response changes a nonlinear gradient.
Exact marginal value Δj(S)
Quadratic training score hj
Surrogate
Ranks update differences using a quadratic model of probe loss at the shared checkpoint.
Appendix
Table 4: Correspondence between the linear analysis and TATC.
Interpreting properties of LLM-generated data is important for understanding its utility and limitations across learning tasks. In this work, we characterize synthetic data through sample-level learnability, studying variation among LLM families and scales, alongside human-written data as a reference. We first generate synthetic datasets spanning single- and multi-label classification, labeling, and tree prediction tasks. We then derive empirical data distributions from encoder training dynamics for both machine and organic data, and estimate the robustness of these distributions across encoders. Finally, we evaluate how data selection strategies based on these learnability signals affect both data sources differently.
Irene Lago, Ana Ezquerro, David Vilares
Universidade da Coruña, CITIC, Spain · Graz University of Technology, IML, Austria
Synthetic data is useful only when the added samples fill missing parts of the training distribution that matter for the downstream task. We introduce LiBaGS, a lightweight, generator-agnostic method for targeted synthetic training data selection. LiBaGS scores candidate synthetic samples by combining decision-boundary proximity, predictive uncertainty, real-data density, and support validity, so that selected samples are both informative and likely to remain on the real data manifold. We then use a boundary-gap allocation rule that targets sparse but realistic decision-boundary neighborhoods, rather than simply adding more data or selecting only the most uncertain candidates. LiBaGS also learns when enough synthetic samples have been added through a marginal-value stopping rule, assigns softer labels near ambiguous boundaries, and uses a diversity objective to avoid redundant near-duplicate selections. Experiments show that LiBaGS improves accuracy over classical oversampling, hard augmentation, uncertainty and density ablations, and targeted-generation selection criteria.
Abhishek Moturu, Anna Goldenberg, Babak Taati
Department of Computer Science University of Toronto The Hospital for Sick Children UHN KITE Research Institute T-CAIREM Vector Institute · Department of Computer Science Department of Laboratory Medicine and Pathobiology University of Toronto The Hospital for Sick Children T-CAIREM Vector Institute · Department of Computer Science Institute of Biomedical Engineering University of Toronto Rehabilitation Sciences Institute UHN KITE Research Institute Vector Institute
Synthetic data now is essential to LLM training, used to strengthen advanced capabilities such as autonomous and long-horizon task execution. Yet recent work shows that training on it at scale can degrade model generation, making it important to decide what synthetic data is worth training on. While current data curation practices do so with individual-level signals (i.e., estimates of each data sample's training utility in isolation), across pre-training and post-training settings we show that this is insufficient for synthetic data, and that group-level signals (i.e., estimates of utility that account for interactions among data samples) are necessary for effective data curation. First, we show that individual-level signals are blind to how samples jointly affect training: synthetic datasets with different compositions can be indistinguishable under individual-level influence yet differ sharply under group-level influence, and curating by the latter yields better downstream performance, particularly in generative capability. Second, we find that group-level signals matter more as training pipelines become increasingly synthetic: among widely used data curation methods, only those incorporating them improve over baseline, with gains increasing when weights capturing relations among samples are amplified. Finally, we translate these findings into practice -- for model developers under a compute budget, we offer a cheap diagnostic that prioritizes which groups of synthetic data most need group-level estimation, recovering much of the benefit of full group-level scoring at a fraction of the compute cost.
Cathy Jiao, Chenyan Xiong
Language Technologies Institute, Carnegie Mellon University