Interpreting properties of LLM-generated data is important for understanding its utility and limitations across learning tasks. In this work, we characterize synthetic data through sample-level learnability, studying variation among LLM families and scales, alongside human-written data as a reference. We first generate synthetic datasets spanning single- and multi-label classification, labeling, and tree prediction tasks. We then derive empirical data distributions from encoder training dynamics for both machine and organic data, and estimate the robustness of these distributions across encoders. Finally, we evaluate how data selection strategies based on these learnability signals affect both data sources differently.
Figures & tables
Figure 1: Human ( H ) vs. synthetic training dynamics with Gemma 27B ( G27 ) and the Difficulty-aware prompt on SST-5 (top), EC (middle-top) and PoS tagging (middle-bottom) and dependency parsing based on EWT (bottom). For display, we subsampled the EWT dataset to 20 000 uniformly sampled tokens.
Figure 2: Data maps on SA (SST-5) with LLM-generated samples with the Difficulty-aware prompt. We include: Gemma 4B ( G4 ) Gemma 12B ( G12 ) and 27B ( G27 ), Llama ( L ), Olmo ( O ) and Qwen ( Q ).
Figure 3: Comparison of training dynamics across prompting strategies. Subfigures illustrate the Cliff’s delta for confidence ( γ ) and variability ( σ ) across SST-5, EC, EWT (PoS) and EWT (DP) relatively to the Naive prompt. Colors denote the prompt strategy ( ∙ Exemplified , ∙ Difficulty-aware , ∙ Ambiguous targeted ), while symbols correspond to the LLM used for data generation ( ◊G4 , ◯G12 , □G27 , △L , ▽O ).
Figure 4: Comparison of human and synthetic distributions across LLMs and the Difficulty-aware prompt using Wasserstein distance on T∗θ (lower triangular matrices), and Cliff’s delta, evaluated on correctness (upper triangular matrices). Same notation as in Figure 2 .
Figure 5: Differences on the training dynamics across encoders with synthetic (top) and human (bottom) data on SST-5. Subfigures show instance-level differences in confidence, variability, and correctness relative to the original RoBERTa encoder used in § 5.1.1 . Here, r denotes the percentage of ambiguous samples (top 33% highest variability) shared with RoBERTa.
in-distribution
out-of-distribution
SST-5
EC
EWT PoS
EWT DP
SST-5
EWT PoS
EWT DP
full
L
40.7
38.6
×
×
38.9
×
×
O
43.0
26.4
71.9
×
37.6
74.9
×
Q
40.4
40.8
82.2
8.5
38.7
86.7
8.6
G12
38.3
38.4
80.5
24.9
44.1
86.6
27.9
H
58.1
59.0
93.0
93.2
57.1
94.4
92.3
Table 1: Performance on full data and sample subsets of LLMs selected via uniform sampling: uniform selection ( random ), maximizing ( max ) or minimizing ( min ) confidence ( γ ), correctness ( α ) and variability ( σ ); and approximating the variability threshold ( mid. α ). The last row ( H ) corresponds to the performance with the human-counterpart dataset. Colors are independently normalized per column on a continuous scale from low (red) to high (green) performance to visually illustrate trends. RoBERTa-large is used both to compute α , γ , and σ ; and to train on the selected subsets.
Appendix figures & tables21 assets
Supplementary material from the paper’s appendix.
Appendix
SA
PoS/DP
temperature
top- p
top- k
temperature
top- p
top- k
L
0.8
0.95
50
×
×
×
O
0.8
0.95
50
0.5
0.9
20
Q
1.2
0.95
50
1
0.9
20
G4
1
0.95
50
0.7
0.9
20
G12
1
0.95
50
0.8
0.9
20
Appendix
Table 2: LLM configuration for prompting.
Figure 6: Examples from our synthetic datasets.
Figure 7: Prompt templates. Note that the exemplified (Figure 7(b) ) and Difficulty-aware (Figure 7(c) ) extend the naive prompt (Figure 7(a) ); and the ambiguous-targeted (Figure 7(d) ) extends the Difficulty-aware prompt.
BLEU
ROUGE
TER
METEOR
COMET
Spanish
43.2
0.77
29.8
0.77
0.83
French
38.0
0.72
35.1
0.73
0.82
Appendix
Table 3: Translation quality for SST-5 data.
Figure 8: Differences on the training dynamics across encoders with synthetic (top) and human (bottom) data on EC. Same notation as in Figure 5 .
Figure 9: Differences on the training dynamics across encoders with synthetic (top) and human (bottom) data on PoS tagging. Same notation as in Figure 5 .
Figure 10: Differences on the training dynamics across encoders with synthetic (top) and human (bottom) data on dependency parsing. Same notation as in Figure 5 .
Figure 11: Dataset cartographies on dependency parsing task with LLM-generated English samples with the Difficulty-aware prompt and RoBERTa encoder. The first column corresponds to the synthetic datasets after applying the corrections, whereas the second column contains only the samples that originally preserved the tree structure. Both distributions were subsampled to 20,000 instances to facilitate visualization. Same notation as in Figure 4 .
SST-5
EC
EWT PoS
EWT DP
W1
α
σ
γ
W1
α
σ
γ
W1
α
σ
γ
W1
α
σ
γ
exemp.
L
0.11
0.22
0.06
0.39
0.03
-0.08
-0.13
-0.02
×
×
×
×
×
×
×
×
O
0.09
0.18
-0.08
0.32
0.10
-0.26
0.12
-0.25
0.18
0.07
0.13
0.02
×
×
×
×
Q
0.03
0.03
0.11
0.06
0.03
0.10
-0.01
0.12
0.07
0.04
0.01
0.07
0.03
0.09
-0.06
0.11
G4
0.06
0.12
-0.08
0.21
0.03
0.11
0.04
0.12
0.03
-0.11
-0.03
-0.17
×
×
×
×
G12
0.01
-0.02
0.05
-0.06
0.01
0.06
0.04
0.07
0.10
-0.06
0.07
-0.12
0.08
0.13
-0.20
0.32
Appendix
Table 4: Comparison of prompt strategies with respect to the naive prompt using the Wasserstein distance ( W1 ) and Cliff’s delta ( δ ) over confidence ( γ ), variability ( σ ) and correctness ( α ) across the English datasets. Same notation as Figure 4 for LLM symbols. The symbol ( × ) denotes runs where the LLM could not produce valid formatted or meaningful samples.
in-distribution
out-of-distribution
SST-5
EC
EWT PoS
EWT DP
SST-5
EWT PoS
EWT DP
naive
L
40.5
30.9
×
×
49.2
×
×
O
35.6
27.4
75.0
×
46.1
77.9
×
Q
37.5
39.7
81.1
12.1
46.0
86.2
9.8
G4
39.3
33.7
69.5
×
44.4
73.0
×
G12
42.2
35.3
81.0
12.7
47.6
86.4
12.0
Appendix
Table 5: Performance of RoBERTa on synthetic data generated from different prompts, LLMs and English datasets (SST-5 for sentiment analysis, EC for multi-label classification and EWT for PoS-tagging and dependency parsing). The symbol ( × ) denotes runs where the LLM could not produce valid formatted or meaningful samples. Last row ( H ) denotes the performance with the full original dataset.
Figure 12: Comparison of training dynamics across prompting strategies on Spanish datasets. Subfigures illustrate the Cliff’s delta for confidence ( γ ) and variability ( σ ) across SST-5, EC, EWT (PoS) and EWT (DP) relatively to the naive prompt. Same notation as in Figure 3 .
Figure 13: Comparison of human and synthetic Spanish distributions across LLMs and the Difficulty-aware prompt using Wasserstein distance on T∗θ (lower triangular matrices), and Cliff’s delta on correctness (upper triangular matrices). Same notation as in Figure 2 .
Figure 14: Differences on the training dynamics across encoders with synthetic (top) and human (bottom) data on Spanish SST-5. Same notation as in Figure 5 .
Figure 15: Comparison of training dynamics across prompting strategies on French datasets. Subfigures illustrate the Cliff’s delta for confidence ( γ ) and variability ( σ ) across SST-5, EC, EWT (PoS) and EWT (DP) relatively to the naive prompt. Same notation as in Figure 3 .
Figure 16: Comparison of human and synthetic French distributions across LLMs and the Difficulty-aware prompt using Wasserstein distance on T∗θ (lower triangular matrices), and Cliff’s delta on correctness (upper triangular matrices). Same notation as in Figure 2 .
Figure 17: Differences on the training dynamics across encoders with synthetic (top) and human (bottom) data on French SST-5. Same notation as in Figure 5 .
in-distribution
out-of-distribution
SST-5
EC
EWT PoS
EWT DP
SST-5
EWT PoS
EWT DP
es
fr
es
es
fr
es
fr
es
fr
es
fr
es
fr
naive
L
37.6
35.4
4.9
×
×
×
×
35.3
42.6
×
×
×
×
O
34.4
36.2
24.2
75.5
68.5
×
×
45.9
36.2
79.5
71.0
×
×
Q
36.3
37.6
36.0
87.5
84.9
7.2
7.3
42.8
46.4
89.5
87.7
9.5
6.8
G4
32.1
32.9
30.9
77.6
77.9
×
×
43.2
44.5
80.3
80.4
×
×
Appendix
Table 6: Performance of XLM-RoBERTa on synthetic data, generated with different prompts and LLMs in our multilingual benchmark. Same notation as in Table 5 .
in-distribution
out-of-distribution
SST-5
EC
EWT PoS
EWT DP
SST-5
EWT PoS
EWT DP
es
fr
es
es
fr
es
fr
es
fr
es
fr
es
fr
full
L
23.1
35.7
32.6
×
×
×
×
20.0
39.7
×
×
×
×
O
38.7
39.4
23.5
57.7
56.8
×
×
41.4
40.1
58.2
58.8
×
×
Q
37.2
38.9
38.1
87.8
85.8
8.2
8.3
45.0
45.8
88.9
87.9
10.6
8.0
G12
40.2
37.5
35.0
87.6
88.3
23.7
28.3
46.0
44.5
89.0
89.5
26.9
30.5
Appendix
Table 7: Performance comparison on sample subsets for Spanish and French datasets. Same criteria and notation as in Table 1 . We use ISO-639 codes to denote languages.
Figure 18: Dataset cartographies on all tasks with human and LLM-generated English samples with the Difficulty-aware prompt and RoBERTa encoder. Same notation as in Figure 4 .
Figure 19: Dataset cartographies on all tasks with human and LLM-generated Spanish samples with the Difficulty-aware prompt and RoBERTa encoder. Same notation as in Figure 4 .
Figure 20: Dataset cartographies on all tasks with human and LLM-generated French samples with the Difficulty-aware prompt and RoBERTa encoder. Same notation as in Figure 4 .
Synthetic data now is essential to LLM training, used to strengthen advanced capabilities such as autonomous and long-horizon task execution. Yet recent work shows that training on it at scale can degrade model generation, making it important to decide what synthetic data is worth training on. While current data curation practices do so with individual-level signals (i.e., estimates of each data sample's training utility in isolation), across pre-training and post-training settings we show that this is insufficient for synthetic data, and that group-level signals (i.e., estimates of utility that account for interactions among data samples) are necessary for effective data curation. First, we show that individual-level signals are blind to how samples jointly affect training: synthetic datasets with different compositions can be indistinguishable under individual-level influence yet differ sharply under group-level influence, and curating by the latter yields better downstream performance, particularly in generative capability. Second, we find that group-level signals matter more as training pipelines become increasingly synthetic: among widely used data curation methods, only those incorporating them improve over baseline, with gains increasing when weights capturing relations among samples are amplified. Finally, we translate these findings into practice -- for model developers under a compute budget, we offer a cheap diagnostic that prioritizes which groups of synthetic data most need group-level estimation, recovering much of the benefit of full group-level scoring at a fraction of the compute cost.
Cathy Jiao, Chenyan Xiong
Language Technologies Institute, Carnegie Mellon University
Synthetic data are increasingly used to scale LLM training, yet more synthetic data do not necessarily produce better models. Useful synthetic data must add information relevant to the target task without introducing errors that offset their benefit, and the value of an example can change as the training set grows. We develop a linear theory that characterizes this tradeoff and determines where synthetic data are useful, how much should be added, and the marginal value of adding one example to an existing set. The analysis shows the conditions when input coverage alone is sufficient and when synthetic errors must also be considered. Guided by these results, we introduce \emph{Training-Aware Target Coverage} (TATC), a synthetic data selection method for LLM fine-tuning. TATC identifies candidates whose training effects are beneficial to the target task and selects among them to expand coverage of target-relevant directions not already represented by the available data. Experiments on text and image data verify the linear theory. With mathematical reasoning tasks, TATC selects synthetic solutions for fine-tuning Qwen2.5-Math-1.5B-Instruct and outperforms alternative synthetic-data selection methods on GSM8K across selection budgets. In summary, we provide a principled approach to synthetic data selection by quantifying and maximizing its value to the target task.
Yang Ba, Michelle V. Mancenido, Rong Pan
School of Computing and Augmented Intelligence, Arizona State University · School of Mathematical and Natural Sciences, Arizona State University
Synthetic data becomes crucial for large language model training, but its effectiveness is highly inconsistent. We provide an information-theoretic account of this inconsistency: synthetic data improves a model only when the generation-training loop is information-open, i.e., shaped by external signals (verifiers, environments, or rubrics) that inject task-relevant information beyond the model's current distribution. When the loop is information-closed (relying on the model's own outputs without such signals), the data processing inequality ensures that task-relevant information can only decrease, making collapse a predicted outcome. Among information-open pipelines, both efficiency and generalization hinge on the meta-level of supervision: a coarser signal such as binary correctness treats all acceptable outputs as equivalent, so the behavior it teaches is not tied to any particular domain or surface form and generalizes naturally across tasks and domains. These observations lead to a guiding thesis: learning preferentially converges to the most information-efficient signal component available, which accelerates learning when that component is the intended one, but causes reward hacking when a spurious pattern happens to be simpler.
Hanyu Li, Zhengqi Sun, Xiaotie Deng
CFCS, School of Computer Science, Peking University, Beijing, China · Department of Information Management, Peking University, Beijing, China