Interpreting properties of LLM-generated data is important for understanding its utility and limitations across learning tasks. In this work, we characterize synthetic data through sample-level learnability, studying variation among LLM families and scales, alongside human-written data as a reference. We first generate synthetic datasets spanning single- and multi-label classification, labeling, and tree prediction tasks. We then derive empirical data distributions from encoder training dynamics for both machine and organic data, and estimate the robustness of these distributions across encoders. Finally, we evaluate how data selection strategies based on these learnability signals affect both data sources differently.
Figures & tables
Figure 1: Human ( H ) vs. synthetic training dynamics with Gemma 27B ( G27 ) and the Difficulty-aware prompt on SST-5 (top), EC (middle-top) and PoS tagging (middle-bottom) and dependency parsing based on EWT (bottom). For display, we subsampled the EWT dataset to 20 000 uniformly sampled tokens.
Figure 2: Data maps on SA (SST-5) with LLM-generated samples with the Difficulty-aware prompt. We include: Gemma 4B ( G4 ) Gemma 12B ( G12 ) and 27B ( G27 ), Llama ( L ), Olmo ( O ) and Qwen ( Q ).
Figure 3: Comparison of training dynamics across prompting strategies. Subfigures illustrate the Cliff’s delta for confidence ( γ ) and variability ( σ ) across SST-5, EC, EWT (PoS) and EWT (DP) relatively to the Naive prompt. Colors denote the prompt strategy ( ∙ Exemplified , ∙ Difficulty-aware , ∙ Ambiguous targeted ), while symbols correspond to the LLM used for data generation ( ◊G4 , ◯G12 , □G27 , △L , ▽O ).
Figure 4: Comparison of human and synthetic distributions across LLMs and the Difficulty-aware prompt using Wasserstein distance on T∗θ (lower triangular matrices), and Cliff’s delta, evaluated on correctness (upper triangular matrices). Same notation as in Figure 2 .
Figure 5: Differences on the training dynamics across encoders with synthetic (top) and human (bottom) data on SST-5. Subfigures show instance-level differences in confidence, variability, and correctness relative to the original RoBERTa encoder used in § 5.1.1 . Here, r denotes the percentage of ambiguous samples (top 33% highest variability) shared with RoBERTa.
in-distribution
out-of-distribution
SST-5
EC
EWT PoS
EWT DP
SST-5
EWT PoS
EWT DP
full
L
40.7
38.6
×
×
38.9
×
×
O
43.0
26.4
71.9
×
37.6
74.9
×
Q
40.4
40.8
82.2
8.5
38.7
86.7
8.6
G12
38.3
38.4
80.5
24.9
44.1
86.6
27.9
H
58.1
59.0
93.0
93.2
57.1
94.4
92.3
Table 1: Performance on full data and sample subsets of LLMs selected via uniform sampling: uniform selection ( random ), maximizing ( max ) or minimizing ( min ) confidence ( γ ), correctness ( α ) and variability ( σ ); and approximating the variability threshold ( mid. α ). The last row ( H ) corresponds to the performance with the human-counterpart dataset. Colors are independently normalized per column on a continuous scale from low (red) to high (green) performance to visually illustrate trends. RoBERTa-large is used both to compute α , γ , and σ ; and to train on the selected subsets.
Appendix figures & tables21 assets
Supplementary material from the paper’s appendix.
Appendix
SA
PoS/DP
temperature
top- p
top- k
temperature
top- p
top- k
L
0.8
0.95
50
×
×
×
O
0.8
0.95
50
0.5
0.9
20
Q
1.2
0.95
50
1
0.9
20
G4
1
0.95
50
0.7
0.9
20
G12
1
0.95
50
0.8
0.9
20
Appendix
Table 2: LLM configuration for prompting.
Figure 6: Examples from our synthetic datasets.
Figure 7: Prompt templates. Note that the exemplified (Figure 7(b) ) and Difficulty-aware (Figure 7(c) ) extend the naive prompt (Figure 7(a) ); and the ambiguous-targeted (Figure 7(d) ) extends the Difficulty-aware prompt.
BLEU
ROUGE
TER
METEOR
COMET
Spanish
43.2
0.77
29.8
0.77
0.83
French
38.0
0.72
35.1
0.73
0.82
Appendix
Table 3: Translation quality for SST-5 data.
Figure 8: Differences on the training dynamics across encoders with synthetic (top) and human (bottom) data on EC. Same notation as in Figure 5 .
Figure 9: Differences on the training dynamics across encoders with synthetic (top) and human (bottom) data on PoS tagging. Same notation as in Figure 5 .
Figure 10: Differences on the training dynamics across encoders with synthetic (top) and human (bottom) data on dependency parsing. Same notation as in Figure 5 .
Figure 11: Dataset cartographies on dependency parsing task with LLM-generated English samples with the Difficulty-aware prompt and RoBERTa encoder. The first column corresponds to the synthetic datasets after applying the corrections, whereas the second column contains only the samples that originally preserved the tree structure. Both distributions were subsampled to 20,000 instances to facilitate visualization. Same notation as in Figure 4 .
SST-5
EC
EWT PoS
EWT DP
W1
α
σ
γ
W1
α
σ
γ
W1
α
σ
γ
W1
α
σ
γ
exemp.
L
0.11
0.22
0.06
0.39
0.03
-0.08
-0.13
-0.02
×
×
×
×
×
×
×
×
O
0.09
0.18
-0.08
0.32
0.10
-0.26
0.12
-0.25
0.18
0.07
0.13
0.02
×
×
×
×
Q
0.03
0.03
0.11
0.06
0.03
0.10
-0.01
0.12
0.07
0.04
0.01
0.07
0.03
0.09
-0.06
0.11
G4
0.06
0.12
-0.08
0.21
0.03
0.11
0.04
0.12
0.03
-0.11
-0.03
-0.17
×
×
×
×
G12
0.01
-0.02
0.05
-0.06
0.01
0.06
0.04
0.07
0.10
-0.06
0.07
-0.12
0.08
0.13
-0.20
0.32
Appendix
Table 4: Comparison of prompt strategies with respect to the naive prompt using the Wasserstein distance ( W1 ) and Cliff’s delta ( δ ) over confidence ( γ ), variability ( σ ) and correctness ( α ) across the English datasets. Same notation as Figure 4 for LLM symbols. The symbol ( × ) denotes runs where the LLM could not produce valid formatted or meaningful samples.
in-distribution
out-of-distribution
SST-5
EC
EWT PoS
EWT DP
SST-5
EWT PoS
EWT DP
naive
L
40.5
30.9
×
×
49.2
×
×
O
35.6
27.4
75.0
×
46.1
77.9
×
Q
37.5
39.7
81.1
12.1
46.0
86.2
9.8
G4
39.3
33.7
69.5
×
44.4
73.0
×
G12
42.2
35.3
81.0
12.7
47.6
86.4
12.0
Appendix
Table 5: Performance of RoBERTa on synthetic data generated from different prompts, LLMs and English datasets (SST-5 for sentiment analysis, EC for multi-label classification and EWT for PoS-tagging and dependency parsing). The symbol ( × ) denotes runs where the LLM could not produce valid formatted or meaningful samples. Last row ( H ) denotes the performance with the full original dataset.
Figure 12: Comparison of training dynamics across prompting strategies on Spanish datasets. Subfigures illustrate the Cliff’s delta for confidence ( γ ) and variability ( σ ) across SST-5, EC, EWT (PoS) and EWT (DP) relatively to the naive prompt. Same notation as in Figure 3 .
Figure 13: Comparison of human and synthetic Spanish distributions across LLMs and the Difficulty-aware prompt using Wasserstein distance on T∗θ (lower triangular matrices), and Cliff’s delta on correctness (upper triangular matrices). Same notation as in Figure 2 .
Figure 14: Differences on the training dynamics across encoders with synthetic (top) and human (bottom) data on Spanish SST-5. Same notation as in Figure 5 .
Figure 15: Comparison of training dynamics across prompting strategies on French datasets. Subfigures illustrate the Cliff’s delta for confidence ( γ ) and variability ( σ ) across SST-5, EC, EWT (PoS) and EWT (DP) relatively to the naive prompt. Same notation as in Figure 3 .
Figure 16: Comparison of human and synthetic French distributions across LLMs and the Difficulty-aware prompt using Wasserstein distance on T∗θ (lower triangular matrices), and Cliff’s delta on correctness (upper triangular matrices). Same notation as in Figure 2 .
Figure 17: Differences on the training dynamics across encoders with synthetic (top) and human (bottom) data on French SST-5. Same notation as in Figure 5 .
in-distribution
out-of-distribution
SST-5
EC
EWT PoS
EWT DP
SST-5
EWT PoS
EWT DP
es
fr
es
es
fr
es
fr
es
fr
es
fr
es
fr
naive
L
37.6
35.4
4.9
×
×
×
×
35.3
42.6
×
×
×
×
O
34.4
36.2
24.2
75.5
68.5
×
×
45.9
36.2
79.5
71.0
×
×
Q
36.3
37.6
36.0
87.5
84.9
7.2
7.3
42.8
46.4
89.5
87.7
9.5
6.8
G4
32.1
32.9
30.9
77.6
77.9
×
×
43.2
44.5
80.3
80.4
×
×
Appendix
Table 6: Performance of XLM-RoBERTa on synthetic data, generated with different prompts and LLMs in our multilingual benchmark. Same notation as in Table 5 .
in-distribution
out-of-distribution
SST-5
EC
EWT PoS
EWT DP
SST-5
EWT PoS
EWT DP
es
fr
es
es
fr
es
fr
es
fr
es
fr
es
fr
full
L
23.1
35.7
32.6
×
×
×
×
20.0
39.7
×
×
×
×
O
38.7
39.4
23.5
57.7
56.8
×
×
41.4
40.1
58.2
58.8
×
×
Q
37.2
38.9
38.1
87.8
85.8
8.2
8.3
45.0
45.8
88.9
87.9
10.6
8.0
G12
40.2
37.5
35.0
87.6
88.3
23.7
28.3
46.0
44.5
89.0
89.5
26.9
30.5
Appendix
Table 7: Performance comparison on sample subsets for Spanish and French datasets. Same criteria and notation as in Table 1 . We use ISO-639 codes to denote languages.
Figure 18: Dataset cartographies on all tasks with human and LLM-generated English samples with the Difficulty-aware prompt and RoBERTa encoder. Same notation as in Figure 4 .
Figure 19: Dataset cartographies on all tasks with human and LLM-generated Spanish samples with the Difficulty-aware prompt and RoBERTa encoder. Same notation as in Figure 4 .
Figure 20: Dataset cartographies on all tasks with human and LLM-generated French samples with the Difficulty-aware prompt and RoBERTa encoder. Same notation as in Figure 4 .