Data selection is critical for training large language models on massive and heterogeneous corpora. Meta-learning for Training-data Selection offers a principled alternative to heuristic scoring by learning data weights from a target validation objective, but existing methods face a trade-off between fine-grained valuation and transferability to unseen data. A natural solution is to replace per-sample weights with a selection network. However, we find that directly incorporating such a network into existing MTS objectives leads to unstable optimization and poor generalization, caused by weight suppression and persistent reliance on easy-to-learn features. To address these issues, we propose Transferable Example Scoring and Selection (TESS), a scalable data-selection framework built on a Pointwise Value Matching objective (PVM). Experiments on LLM safety and targeted instruction tuning demonstrate strong transfer across datasets, from subsets to full corpora, and from smaller to larger models.
Figures & tables
Figure 1: Simulation of the weight dynamics in Theorem 1 . Curves show the weights wi(t) , with fixed pseudo-labels Δi∗ indicated in the legend. Initial weights are (0.2,0.2,0.3,0.2) , and δ0 is 10−6 .
Figure 2: Training dynamics of ScaleBiO and TESS under different meta-network objectives. With SBO loss, most predicted weights shrink toward zero as accuracy declines. In contrast, PVM loss maintains nonvanishing weight quantiles and high accuracy.
Shortcut ( Cov / ρ )
Method
Early Stage
Middle Stage
Late Stage
Avg. Gap
AccS
AccU
Gap
AccS
AccU
Gap
AccS
AccU
Gap
Strong 0.17 / 0.57
TESS + LϕPVM
98.84
79.38
19.46
99.56
84.62
14.94
99.56
84.97
14.59
16.33
TESS + LϕSBO
97.96
72.44
25.52
90.22
52.26
37.96
84.71
46.31
38.40
33.96
ScaleBiO + LϕSBO
93.42
78.84
14.58
90.76
57.73
33.03
76.00
32.44
43.56
30.39
Weak 0.06 / 0.41
TESS + LϕPVM
98.04
91.47
6.57
98.67
93.51
5.16
98.76
93.60
5.16
5.63
TESS + LϕSBO
96.53
80.71
15.82
87.82
65.78
22.04
84.00
60.53
23.47
20.44
Table 1: Safe–unsafe data classification accuracy under different shortcut strengths. Cov and ρ denote the empirical covariance and Pearson correlation coefficient between the shortcut feature and the label, respectively. Early, middle, and late stages correspond to 20% , 50% , and 80% of meta-network training, respectively. AccS and AccU denote seen and unseen accuracy, and Gap=AccS−AccU denotes the generalization gap.
Setting
Training
Generalization
Dataset
Bench
Random
Task-agnostic Heuristics
Learnable Weighting
GradSafe
Bi-Anchor
SEAL w
SBO w
SEAL ϕ
SBO ϕ
TESS
SEAL ϕ
SBO ϕ
TESS
Target LLM: Llama3-8B-Instruct
Alpaca
DH4
25.00
28.00
49.00
26.75
38.25
12.50
12.50
36.75
8.50
6.50
35.50
HB
15.00
16.00
35.00
13.50
21.00
9.00
7.00
20.50
6.00
4.00
25.00
HEx
6.55
8.97
24.58
6.90
10.69
5.86
3.44
11.38
3.10
3.44
15.52
Table 2: Attack success rate (ASR, %; higher is better) for harmful-data selection on Alpaca and Dolly. Best and second-best results are bold and underlined , ranked separately within two settings.
Setting
Method
Alpaca
Dolly
Avg.
DH4
HB
HEx
DH4
HB
HEx
Training
TESS w/ PVM Loss
44.50
23.50
24.83
86.50
87.00
88.62
59.16
TESS w/ SBO Loss
39.50
19.50
22.06
58.25
40.50
45.52
37.56-21.60
TESS w/ KL Loss
40.75
19.00
23.86
60.75
53.50
59.66
42.92-16.24
Pseudo-Label Ranking
11.25
7.50
5.86
75.75
75.00
70.69
41.01-18.15
Generalization
TESS w/ PVM Loss
38.75
16.00
18.28
84.50
83.50
80.00
53.51
Table 3: Ablation of Different Training Losses for the Meta-network on Qwen2.5-7B-Instruct.
Method
GSM8K
CodeX
Overall Avg.
N =1K
N =5K
N =10K
Avg.
N =5K
N =10K
N =20K
Avg.
Random
13.41
15.92
16.13
15.15
29.05
27.03
27.70
27.93
21.54
LESS
18.80
22.36
22.59
21.25
25.00
26.35
28.37
26.57
23.91 +2.37
RDS+
17.43
21.60
23.57
20.87
28.37
26.35
29.73
28.15
24.51 +2.97
SEAL w
14.78
16.37
17.89
16.35
24.32
25.00
25.68
25.00
20.67 -0.87
SBO w
15.01
17.74
18.65
17.13
26.35
25.68
27.03
26.35
21.74 +0.20
Table 4: Llama-2-7B performance after fine-tuning on N selected Tulu V2 examples.
Appendix figures & tables1 asset
Supplementary material from the paper’s appendix.
Appendix
Setting
Method
Alpaca
Dolly
Overall ASR
DH4
HB
HEx
Srel
DH4
HB
HEx
Srel
Avg.
Retention
Training
Big2Big
44.50
23.50
24.83
1.00 ×
86.50
87.00
88.62
1.00 ×
59.16
–
Small2Big
35.50
13.00
13.79
6.60 ×
74.00
72.00
66.90
5.65 ×
45.87
77.53
Generalization
Big2Big
38.75
16.00
18.28
1.00 ×
84.50
83.50
80.00
1.00 ×
53.51
–
Small2Big
32.25
13.00
14.14
6.20 ×
77.50
70.00
61.72
5.65 ×
44.77
83.67
Appendix
Table 5: Small-to-large model transfer with TESS. Small2Big uses Qwen2.5-0.5B-Instruct for meta-network training and Qwen2.5-7B-Instruct for fine-tuning on the selected data; Big2Big uses Qwen2.5-7B-Instruct for both. We report ASR (%) and training speedup Srel relative to Big2Big within each dataset and setting. Avg. averages the six ASR scores, and Retention (%) is the ratio of Small2Big’s average ASR to Big2Big’s within the same setting. Higher is better for all metrics; better ASR and speedup results are bold .