Synthetic data now is essential to LLM training, used to strengthen advanced capabilities such as autonomous and long-horizon task execution. Yet recent work shows that training on it at scale can degrade model generation, making it important to decide what synthetic data is worth training on. While current data curation practices do so with individual-level signals (i.e., estimates of each data sample's training utility in isolation), across pre-training and post-training settings we show that this is insufficient for synthetic data, and that group-level signals (i.e., estimates of utility that account for interactions among data samples) are necessary for effective data curation. First, we show that individual-level signals are blind to how samples jointly affect training: synthetic datasets with different compositions can be indistinguishable under individual-level influence yet differ sharply under group-level influence, and curating by the latter yields better downstream performance, particularly in generative capability. Second, we find that group-level signals matter more as training pipelines become increasingly synthetic: among widely used data curation methods, only those incorporating them improve over baseline, with gains increasing when weights capturing relations among samples are amplified. Finally, we translate these findings into practice -- for model developers under a compute budget, we offer a cheap diagnostic that prioritizes which groups of synthetic data most need group-level estimation, recovering much of the benefit of full group-level scoring at a fraction of the compute cost.
Figures & tables
Figure 1
Figure 2: Semantic/topic diversity of our candidate pool on the pre-training corpus.
Figure 3: Range of group oracle and individual proxy oracle scores across our constructed groups (on groups of 5120 and 128 samples). See Appendix D.4 for full sweep over group sizes.
Figure 4: (a) Relative improvement over the base model by task type in pre-training. (b) Pass@8 over GRPO training steps in post-training.
Method
Representation
Individual
GradSim ( Pruthi et al., 2020 )
Gradient
FineWeb-Edu ( Penedo et al., 2024 )
Classifier score
BM25 ( Trotman et al., 2014 )
Lexical (tokens)
Group
GREATS ( Wang et al., 2024 )
Gradient
GMRel ( Yu et al., 2025 )
Learned influence
Table 2: Taxonomy of individual-and group-level data selection methods. Details in Appendix B.5
Figure 5: (a, b) Performance of individual-/group- level selection methods as rephrase factor increases. (c) Group-oracle value and diversity trade-off. (d) Increasing GMRel’s relational weight at r=50 .
Figure 6: Gap between the group and individual proxy oracle diverges as within-group diversity increases.
Figure 7: Allocating a limited group-oracle budget with cheap diagnostics. See Appendix D.5 for additional analysis plots.
Appendix figures & tables14 assets
Supplementary material from the paper’s appendix.
Appendix
Parameter
Pre-Training Setting
Post-training Setting
Probe model
Repro-400M (trained on 6.59B Repro tokens ( Yu et al., 2025 ) )
Qwen2.5-1.5B-Instruct
Tokenizer
Pythia-410m
Qwen2.5-1.5B-Instruct
Optimizer
AdamW, 1 step
AdamW, 1 step
Learning rate
2e-5
1e-5
Precision
bf16
bf16
Grad. accumulation
Yes
Yes
Appendix
Table 3: Oracle scoring configurations for the pre-training and post-training settings.
Setting
size
Quantity
maxj∣Iproxy(Gj,θ)−τj∣
ϵ
diversity range
Pretraining
5120
1000
5.0×10−5
5×10−5
0.095 – 0.212
Pretraining
10240
500
5.0×10−5
5×10−5
0.096 – 0.200
Post-training
128
1000
5.0×10−6
5×10−6
0.156 – 0.435
Post-training
512
500
4.8×10−6
5×10−6
0.157 – 0.425
Appendix
Table 4: Diversity-swept candidate pools. Diversity is 1−cos(G) . Every group’s individual proxy oracle lands within ϵ of its assigned band target, so the pools vary in diversity at (near-)fixed individual proxy oracle.
Pre-training
Post-training
Phase 1: Candidate pool construction
Source dataset
Repro-Rephrased-72B
GooseReason
Semantic clusters
20
20
Group sizes s
5120 (small), 10240 (large)
128 (small), 512 (large)
Candidate pool H
1000 (small), 500 (large)
1000 (small), 500 (large)
Phase 2: Data selection
Appendix
Table 5: Task settings for the pre-training and post-training experiments, organized by the three phases in Section 3.2 . Each training setting is run under two configurations: a small -group and a large -group setting.
Phase 1: Synthetic pool construction
Source documents
Organic web text (DCLM-RefinedWeb)
Rephraser
Repro rephraser-4B ( Yu and Xiong, 2025 )
Sampling
temperature 1.0 , top- p0.9
Pool size N
716,800 sequences
Rephrase factor r
{1,10,50}
Unique source docs ( N/r )
716,800/71,680/14,336
Appendix
Table 6: Experimental settings for the data curation experiments, organized by the three phases in Section 4.1 . The pool is held at a fixed size while the rephrase factor r controls the degree of near-duplicate redundancy.
Parameter
Pre-Training Setting
Post-training Setting
Rollout generation (supervision)
Base model
Repro-400M (as in Table 3 )
Qwen2.5-1.5B-Instruct
Candidate pool
Repro-Rephrased-72B
GooseReason
Source tokenizer
Pythia-410m
Qwen2.5-1.5B-Instruct
Reference set
FLAN (128 samples)
GPQA-extend (128 samples)
Candidates per rollout K
10
10
Appendix
Table 7: GMRel relational model configurations: oracle rollout generation (which produces the supervision) and fitting of the relational influence model.
Allocation
B
q
Ratio (Eq. 8 )
Calculated Ratio
Group oracle
H
0
1
1.000
GradDiv full
B
1
≈B/H+κ
0.790
GradDiv q
B
0.5
≈B/H+κq
0.495
GradDiv q
B
0.1
0.259
GradDiv q
B
0.05
0.229
GradDiv q
B
0.01
0.206
Appendix
Table 8: Cumulative GPU memory of each allocation relative to the full group oracle. The last column is the exact ratio of Equation 8 on the pre-training pool of Figure 7(d) ( H=1000 , ∣G∣=5120 , R=128 , B/H=0.2 , κ=0.593 ).
Method
HellaSwag
ARC-C
PIQA
MMLU
CSQA
LAMBADA
SQuAD
CoQA
Avg
Rephrase factor r=1
Random
30.95
26.96
66.16
23.19
19.41
41.04
9.82
25.52
30.38
GradSim
30.81
26.28
65.34
22.98
21.21
36.37
3.65
24.69
28.92
FineWeb-Edu
30.89
25.51
64.53
22.87
19.16
38.09
17.79
23.60
30.31
BM25
31.14
26.45
65.34
23.21
19.66
38.15
5.40
26.68
29.50
GMRel ( τ≈1.01 )
30.41
25.94
65.02
23.15
19.82
36.76
25.17
24.62
31.36
Appendix
Table 9: Per-task accuracy at the final step (200 steps) for the Repro-400M continuation, by selection method and rephrase factor r . Best average within each r is in bold , and second best is underlined .
Figure 8: Discrepancy between the group oracle Iorcl(G,θ) and individual proxy oracle Iproxy(Gj,θ) scores across varying sizes of both the synthetic pretraining and the RLVR data groups.
Figure 9: Range of group oracle and individual proxy oracle scores across our constructed groups, for each group size in the pre-training and post-training settings.
Figure 10: Topic composition of the k=20 semantic clusters used for proxy-matched diversity sampling. Each slice is one cluster, labeled with its topic and share of the corpus, and colored by topic family.
Figure 11: Diversity metrics on our constructed groups versus randomly formed groups in the pre-training setting.
Figure 12: Diversity metrics on our constructed groups versus randomly formed data groups in the post-training setting.
Figure 13: Choice of diagnostic for group-oracle budget allocation, for pre-train and post-train data (group sizes of 5120 and 512 samples, respectively).
Figure 14: Effect of gradient subsampling on group-oracle budget allocation, comparing the full GradDiv diagnostic against subsampled variants ( q=0.1 , q=0.01 ) for pre-train and post-train data (group sizes of 5120 and 512 samples, respectively).