Training-Aware Target Coverage for Synthetic Data Selection
Organizations: School of Computing and Augmented Intelligence, Arizona State University · School of Mathematical and Natural Sciences, Arizona State University
Abstract
Synthetic data are increasingly used to scale LLM training, yet more synthetic data do not necessarily produce better models. Useful synthetic data must add information relevant to the target task without introducing errors that offset their benefit, and the value of an example can change as the training set grows. We develop a linear theory that characterizes this tradeoff and determines where synthetic data are useful, how much should be added, and the marginal value of adding one example to an existing set. The analysis shows the conditions when input coverage alone is sufficient and when synthetic errors must also be considered. Guided by these results, we introduce \emph{Training-Aware Target Coverage} (TATC), a synthetic data selection method for LLM fine-tuning. TATC identifies candidates whose training effects are beneficial to the target task and selects among them to expand coverage of target-relevant directions not already represented by the available data. Experiments on text and image data verify the linear theory. With mathematical reasoning tasks, TATC selects synthetic solutions for fine-tuning Qwen2.5-Math-1.5B-Instruct and outperforms alternative synthetic-data selection methods on GSM8K across selection budgets. In summary, we provide a principled approach to synthetic data selection by quantifying and maximizing its value to the target task.
Figures & tables
Appendix figures & tables4 assets
Supplementary material from the paper’s appendix.
Appendix
| Partition | Examples | Role |
|---|---|---|
| 1,500 | Fit PCA and the reference linear target | |
| 500 | Estimate the residual variance scale | |
| 1,200 | Define the real training information | |
| 1,000 | Estimate the target moment | |
| 3,800 | Supply candidate inputs |
| Score | Metric | Value |
|---|---|---|
| Coverage | AUROC | 0.519 |
| Complete marginal | AUROC | 0.99 |
| Complete marginal | Correlation | 0.99 |
| Complete marginal | Sign agreement | 99.2% |
| Label model | Label accuracy | Predicted ratio | Test optimum ratio |
|---|---|---|---|
| 12,000 training examples | 0.889 | 7.0 | 8.0 |
| 400 training examples | 0.836 | 0.90 | 0.95 |
| 12 coordinates removed | 0.604 | 0 | 0 |
| Linear quantity or role | TATC quantity | Status | Correct interpretation |
|---|---|---|---|
| Real information | Approximation | Summarizes prompt directions already represented by real data; is numerical stabilization. | |
| Target moment | Sample estimate | Estimates target relevance only in the checkpoint prompt representation. | |
| Selected information | Exact in representation | Each selected input adds a rank-one prompt-feature contribution. | |
| Clean coverage marginal | Exact in representation | Exact decrease in the downstream-weighted inverse-trace criterion. | |
| Systematic label error | Response-dependent gradient and AdamW update | Approximation | For squared loss, response error changes the gradient by ; in an LLM the full response changes a nonlinear gradient. |
| Exact marginal value | Quadratic training score | Surrogate | Ranks update differences using a quadratic model of probe loss at the shared checkpoint. |