Beyond Prompt Count: How Data Shapes Transfer in On-Policy Distillation
Authors: Jiaxuan Wang, Jiafei Lyu, Yuchen Cai, Siye Wu, Pengyuan Wang, Jiashun Liu, Xiang Cheng, Kai Yang, +3 more
Organizations: State Key Laboratory of Novel Software Technology, Nanjing University · School of Intelligence Science and Technology, Nanjing University · LLM Department, Tencent · RUC
On-policy distillation (OPD) trains students using teacher feedback on their own sampled responses, yet how prompt choice shapes transfer across teacher-student pairs remains poorly understood. We systematically study prompt quantity, source, and selection across RL- and SFT-continuation pairs and cross-model settings. We find that OPD can be highly prompt-efficient: a few prompts can approach large-pool performance, with four DAPO prompts matching the observed mathematics score of 3,840 DeepMath prompts. However, prompt utility is relational rather than intrinsic: changing only the teacher can reverse the relative effectiveness of mathematics and code prompts. To characterize these transfer differences, we analyze parameter and functional changes across prompt supports and model pairs. Functional alignment with the teacher varies across supports and target tasks; in continuation pairs, teacher-aligned prediction changes can coexist with weak parameter alignment. Continued OPD on effective supports can restore performance after unfavorable transfer. Finally, targeted selection does not consistently outperform uniform random sampling, and filtering out a source that performs poorly alone yields no consistent gain across three paired support draws. Overall, our results distinguish prompt efficiency from prompt interchangeability and show that effective data choice depends on the teacher-student pair and target capability, with random sampling providing a competitive baseline in the studied settings.
Figures & tables
Figure 1: Prompt efficiency does not imply prompt interchangeability. A single prompt can induce substantial transfer, but outcomes vary with prompt identity and composition. Moreover, a few prompts from one source can match the observed mathematics score of thousands from another. These examples motivate our study of prompt count, source, teacher–student pairing, and selection. In (b), P1 – P3 denote the individual code prompts described in Appendix C.4 , while P1+P2 denotes joint OPD using the first two prompts.
Table 1: Teacher–student pairs and OPD sources. The first three teachers are trained by us; the others are publicly released models. We list constituent data sources; support construction, including mixtures and source filtering, is detailed in Appendix B.3 .
Figure 2: Prompt-count scaling across six teacher–student pairs. Stars mark full-pool OPD; S/T denote the initial student and teacher. Circled points in (d) attain the same observed Math score with 4 DAPO versus 3,840 DeepMath prompts. Training budgets are comparison-specific; full results appear in Appendix C .
Figure 3: Data-source preferences across teacher–student pairs. (a) Source sensitivity with a 30B Instruct teacher. (b) Preference reversals between Math-RL and Code-RL teachers with the same initial student. (c) Cross-domain transfer from the Nemotron teacher. Avg is the unweighted mean of the four displayed benchmark groups; side annotations report differences from DeepMath. Darker and lighter shading mark the highest and second-highest OPD scores within each column.
Table 3: Data-source preferences with a multi-domain RL teacher. Qwen3-4B-Inst-Mix → Qwen3-4B-Instruct-2507; M=48 . Teal marks the best and second-best OPD scores; mixture construction is detailed in Appendix B.3 .
Figure 4: Parameter and functional alignment. (a) Parameter and all-domain functional cosines for 43 checkpoints; each line joins the same endpoint. (b) Teacher-direction projection α versus functional cosine on Math probes, grouped by model pair. (c) Functional cosine versus accuracy: 30B-to-4B Math (left) and Mix-RL-to-4B GPQA (right). Definitions: Appendix E.1 .
Figure 5: GPQA accuracy and output-length changes. All supports contain M=48 prompts; changes are relative to each pair’s initial student. Inset: mean output tokens.
Figure 6: Redirecting unfavorable transfer. (a) CKA relative to full-pool OPD. (b) Bidirectional hidden-state replacement. (c) Performance recovery through continued distillation; vertical ticks mark native references. (d) Functional versus parameter alignment to native endpoints. Reference checkpoints and unequal Code update budgets are specified in Appendix E.5 .
Figure 7: Continuation data shapes recovery at matched update counts. 30B Instruct → Qwen3-4B, with 15 updates per stage and a fresh Adam optimizer in every second-stage run; token budgets are not matched. Orange and blue indicate stage-2 GSM8K and DeepMath, respectively. (a) AIME24/25 validation mean@2; solid/dashed lines start from GSM8K/DeepMath checkpoints. Gray step-0 markers show the shared starting scores of 44.17 / 50.83 . (b) Formal Avg 4 (%), with the y-axis truncated at 30% , grouped by stage-1 source; annotations show the gains from using DeepMath rather than GSM8K in stage 2.
Method
Math
GPQA-D
HE+
LCB
Avg 4
GPU ⋅ h
M=8,C=384
Uniform random
31.09 ± 0.50
39.44 ± 1.24
62.40 ± 0.32
17.67 ± 0.86
37.65 ± 0.38
0
Stratified random
30.36 ± 0.39
38.19 ± 0.58
61.74 ± 0.46
17.93 ± 0.41
37.06 ± 0.25
1.1
Semantic diversity
30.21
39.46
62.88
17.36
37.48
<0.1
Hard selection
28.96
38.07
63.57
18.21
37.20
1.1
Shortest
28.75
39.08
61.28
16.86
36.49
2.2
Table 4: Prompt selection within DeepMath. JustRL-1.5B → DS-1.5B. Random selectors report mean ± sample SD over three sampled supports. Bold marks the column best within each M . Costs estimate selection-only GPU hours; 0 denotes no GPU screening.
Figure 12
Figure 9: Source composition and a four-prompt control. 30B Instruct → Qwen3-4B, 25 updates. DeepMath/GSM8K composition sweep (circles, M=48 ) and DeepMath4-only control (diamonds, M=4 ). Within each seed, the control reuses exactly the four DeepMath prompts in the corresponding 4/44 mixture. Error bars and bands show sample SD across selection seeds 42/43/44.
Appendix figures & tables24 assets
Supplementary material from the paper’s appendix.
Appendix
Parameter
Math / OOD
Code (4B)
Direct-code
Samples per question ( k )
16 / 8
8
16
Temperature
1.0
1.0
0.6
Top- p
1.0
1.0
0.95
Top- k / min- p
−1 / 0.0
−1 / 0.0
−1 / 0.0
Max. generation length
16,384
16,384
4,096
Appendix
Table 11: Evaluation configurations. Lengths are in tokens; all settings report mean@ k and pass@ k , with k=16 for Math and k=8 for OOD.
Source
M
AIME24
AIME25
HMMT Feb
HMMT Nov
Math
GPQA-D
HE+
LCB
Initial
–
22.29
19.79
11.04
8.75
15.47
42.74
77.06
17.71
Teacher
–
62.71
56.46
34.38
39.79
48.33
52.65
83.92
20.07
DeepMath
1
55.21
54.17
32.29
37.29
44.74
51.70
82.39
29.93
DeepMath
48
60.21
53.12
33.33
38.96
46.41
53.22
83.23
27.21
DeepMath
3,840
61.25
55.83
32.92
37.08
46.77
52.15
85.14
29.79
Appendix
Table 12: Math-RL continuation. Math GRPO-500 to Qwen3-4B ( Figure 2 a ), with 3,840 rollouts per OPD run. Math reports mean@16; other tasks report mean@8.
Source / support
M
HE+
MBPP+
LCB
Code macro
Initial
–
76.60
65.15
18.14
53.30
Teacher
–
86.36
71.03
28.00
61.80
Eurus Code (ID 5038)
1
79.04
68.39
26.43
57.95
Eurus Code
48
86.66
70.14
28.57
61.79
Eurus Code
3,840
86.05
70.90
28.14
61.70
Eurus Code (full)
22,618
86.74
71.06
27.86
61.89
Appendix
Table 13: Code-RL continuation. Code GRPO-300 to Qwen3-4B ( Figure 2 b ), at OPD step 100; all scores are mean@8. The principal M=1 support is ID 5038 ( B in Section 5 ); rows below the separator give additional single-prompt experiments.
Source / reference
M
HE+
MBPP+
LCB
Code macro
Initial: direct-code SFT-554
–
44.32
52.88
11.25
36.15
Teacher: Code GRPO-400
–
58.12
58.47
15.96
44.18
Open-R1 Code
1
58.82
58.74
15.96
44.51
Open-R1 Code
256
59.98
59.51
15.57
45.02
Open-R1 Code
3,840
59.68
59.62
16.61
45.30
Open-R1 Code (full)
6,519
58.73
58.80
15.50
44.34
Appendix
Table 14: Direct-code RL continuation. Code GRPO-400 to Qwen3-1.7B-Base after direct-code SFT at step 554 ( Figure 2 c ). OPD endpoints are at step 50; all scores are mean@16.
Source
M
AIME24
AIME25
HMMT Feb
HMMT Nov
Math
GPQA-D
HE+
LCB
Initial
–
27.50
25.63
12.50
11.25
19.22
35.92
54.34
13.50
Teacher
–
53.96
39.38
21.67
23.75
34.69
40.47
64.56
19.36
DAPO
1
42.71
30.42
16.88
17.08
26.77
38.13
60.98
16.29
DAPO
2
49.58
36.25
20.83
20.83
31.87
38.64
62.27
16.71
DAPO
4
47.71
36.04
22.08
24.38
32.55
41.10
63.87
17.50
DAPO
8
51.04
35.00
21.25
22.71
32.50
40.91
64.10
18.64
Appendix
Table 15: JustRL continuation. JustRL-DeepSeek-1.5B to DeepSeek-R1-Distill-Qwen-1.5B ( Figure 2 d ), at OPD step 100. Math reports mean@16; other tasks report mean@8. Full denotes the available source pool; DeepMath M=8 is the random support in the prompt-count sweep.
Source
M
AIME24
AIME25
HMMT Feb
HMMT Nov
Math
GPQA-D
HE+
LCB
Initial
–
22.29
19.79
11.04
8.75
15.47
42.74
77.13
17.71
Teacher
–
75.21
61.67
42.50
56.25
58.91
54.86
82.39
33.14
DeepMath
1
46.0
46.7
24.6
31.9
37.29
48.3
82.4
24.3
DeepMath
48
56.46
45.63
29.38
39.79
42.81
52.15
82.32
26.79
DeepMath
3,840
53.8
48.3
30.2
38.1
42.60
52.0
83.7
27.0
DAPO
1
49.4
40.4
26.2
31.5
36.88
49.7
83.5
26.3
Appendix
Table 16: Cross-model distillation. Qwen3-30B-A3B-Instruct-2507 to Qwen3-4B ( Figure 2 e ), at OPD step 15. Math reports mean@16; other tasks report mean@8. Full-pool endpoints at other training durations appear in Table C.2 .
Source
M
AIME24
AIME25
HMMT Feb
HMMT Nov
Math
GPQA-D
HE+
LCB
Initial
–
13.33
10.00
5.00
5.63
8.49
29.10
59.83
12.36
Teacher
–
75.21
61.67
42.50
56.25
58.91
54.86
82.39
33.14
DeepMath
1
27.3
24.0
16.0
14.2
20.4
30.1
65.5
16.1
DeepMath
8
33.1
26.7
15.0
15.2
22.5
32.5
65.8
19.4
DeepMath
48
35.0
26.9
14.2
17.1
23.3
36.2
66.7
18.1
DeepMath
3,840
36.0
30.8
14.8
16.7
24.6
32.7
68.9
18.7
Appendix
Table 17: Cross-model distillation. Qwen3-30B-A3B-Instruct-2507 to the official Qwen3-1.7B checkpoint ( Figure 2 f ), not the SFT-554 initialization in Table C.1 . Math reports mean@16; other tasks report mean@8.
Source
Step
AIME24
AIME25
HMMT Feb
HMMT Nov
Math
GPQA-D
HE+
LCB
Math GRPO-500 → Qwen3-4B
DeepMath
25
62.92
54.58
32.29
38.75
47.14
52.78
84.91
29.86
DeepMath
50
64.79
56.04
34.38
41.46
49.17
54.29
84.45
28.07
Qwen3-30B-A3B-Instruct-2507 → Qwen3-4B
DeepMath
25
55.0
48.8
27.3
37.9
42.2
48.4
82.8
25.8
DeepMath
50
56.9
49.0
31.7
38.8
44.1
48.4
83.5
26.6
Appendix
Table 18: Additional full-pool OPD endpoints. Math reports mean@16; other tasks report mean@8. Math-RL runs use batch size 1,024, with 25,600 and 51,200 rollouts at OPD steps 25 and 50, respectively.
Figure 10: Support size and learning dynamics. All runs use the Qwen3-4B Math-RL continuation pair. (a–b) Training-time AIME24 and AIME25 validation accuracy (mean@2). (c) Mean rollout response length, in thousands of tokens. (d) Teacher–student discrepancy, reported as logged mean K1 on a logarithmic scale. These validation traces are distinct from the final mean@16/mean@8 evaluations in the complete result tables.
Figure 11: Teacher and prompt-source effects during Qwen3-4B OPD. Math-T and Code-T denote Math GRPO-500 and Code GRPO-300; Math48 and Code48 denote the fixed DeepMath and Eurus Code supports. (a) AIME25 validation with Math48. (b) TACO validation with Code48. Both report training-log mean@2. (c) Mean rollout response length. (d) Teacher–student discrepancy, reported as logged mean K1 (log scale). Panels (a–b) use different validation tasks; the common-benchmark source-preference reversal is shown separately in Figure 3 b .
Math-RL teacher
Code-RL teacher
Support / training seed
DeepMath48
Code48
DeepMath48
Code48
42 / 42
54.30
46.84
44.94
50.75
43 / 42
54.20
46.62
44.79
51.27
44 / 42
54.23
46.17
45.22
50.51
42 / 43
53.87
46.37
44.69
50.31
42 / 44
54.26
46.66
44.67
50.42
Appendix
Table 19: Seed robustness of teacher-dependent source preferences. Avg 4 (%) across support-selection and training seeds.
Data
Tokens (M)
Math
GPQA
HE+
LCB
GSM8K
1.19
31.15
46.84
81.25
16.79
+ token match
16.14
31.56
48.04
82.32
17.43
DeepMath
16.63
42.81
52.15
82.32
26.79
Appendix
Table 20: Benchmark-level supervision-token control. M=48 ; tokens are cumulative supervised tokens in millions, and GSM8K token matching to DeepMath is approximate.
Figure 12: Additional source-dependent functional diagnostics. (b) 30B-to-4B. (c) Nemotron. (d) Mix-RL. Upper plots report Math functional cosine and teacher-direction coefficient α ; lower matrices report task-specific KL closure (%). Closure uses teacher-first conditional KL on the initial student’s frozen top-128 support, not a task-score recovery rate. Blue and red indicate negative and positive closure, respectively; colors saturate at −100% and 100% .
Model / OPD support
GPQA-D
HumanEval+
LCB v6
(1) Math-RL: Qwen3-4B → Qwen3-4B
Initial (Compact)
42.74 / 870
77.06 / 325
17.71 / 557
Teacher (GRPO-500)
52.65 / 2,619
83.92 / 700
20.07 / 2,670
DeepMath M48 (Compact)
53.22 / 5,590
83.23 / 2,328
27.21 / 9,567
(2) Qwen3-30B-A3B-Instruct-2507 → Qwen3-4B
Initial
42.74 / 870
77.13 / 325
17.71 / 557
Appendix
Table 21: Accuracy and output length for seven teacher–student pairs. Each cell gives mean@8 accuracy (%) / mean output tokens; s denotes the OPD update step.
Figure 13: Code recovery in parameter and function space. PCA views of the initial student ( I ), degraded single-prompt endpoint ( A′ ), recovered endpoint ( R ), and native Code-M48 endpoint ( N ). Native Code-M48 is trained directly from I ; recovery continues from A′ on an effective 48-prompt support. From A′ to R , the reported distance to N increases from 1.46 to 1.69 in parameter space but decreases from 1.03 to 0.29 in function space. Axis percentages report explained variance; arrows connect endpoints rather than tracing intermediate updates.
Figure 14: Training trajectories under single-prompt and two-prompt supports. All runs use the same Qwen3-4B initialization and Code GRPO-300 teacher. A′ uses prompt 15240; M2 jointly uses prompts 15240 and 187 from initialization. (a) Code accuracy: the unweighted mean of mean@8 over HumanEval+, MBPP+, and LCB-v6. Markers for M2 and prompt 187 alone show only their step-100 endpoints. (b) Functional cosine to the teacher’s change. (c) Normalized functional-update magnitude of A′ , ∥vOPD∥/∥vT∥ .
Figure 15: Parameter displacements after bidirectional support switching. Cosines compare FP32 endpoint-minus-original-initialization vectors for the four paths in Figure 7 . Every endpoint has 30 total updates; G denotes GSM8K and D denotes DeepMath. The G→D versus D→D comparison uses the step-30 reference, unlike the step-15 Math reference in Figure 6 (d) .
Selection seed
Mixed
Filtered
Shared prompts
DeepMath / DAPO / GSM8K
DeepMath / DAPO / GSM8K
42
17 / 15 / 16
21 / 27 / 0
32
43
13 / 15 / 20
22 / 26 / 0
28
44
20 / 13 / 15
31 / 17 / 0
33
Appendix
Table 22: Mixed-pool support composition and overlap. Shared counts prompt instances selected by both arms under the same random ordering, not shared sampled responses.
Benchmark
Selection seed
Mixed
Filtered
Δ
Math average
42 / 43 / 44
43.2 / 40.3 / 38.8
41.4 / 39.3 / 40.7
−1.8 / −0.9 / +1.9
Mean ± SD
40.7±2.3
40.5±1.0
−0.3±2.0
AIME24
42 / 43 / 44
55.8 / 54.2 / 49.2
56.5 / 51.3 / 49.8
+0.6 / −2.9 / +0.6
Mean ± SD
53.1±3.5
52.5±3.5
−0.6±2.0
AIME25
42 / 43 / 44
50.4 / 48.1 / 43.8
44.4 / 46.7 / 46.9
−6.0 / −1.5 / +3.1
Mean ± SD
47.4±3.4
46.0±1.4
−1.5±4.6
Appendix
Table 23: Math results for mixed-pool and source-filtered sampling. 30B Instruct → Qwen3-4B, M=48 , 25 optimization steps; selection seeds 42/43/44 share a fixed training seed. Scores are mean@16 (%); Math average is the unweighted mean of the four benchmarks. Δ denotes paired Filtered − Mixed differences; means, SDs, and differences use unrounded scores.
Benchmark
Selection seed
Mixed
Filtered
Δ
OOD average
42 / 43 / 44
55.2 / 54.9 / 53.9
55.2 / 53.7 / 54.6
0.0 / −1.2 / +0.7
Mean ± SD
54.7±0.7
54.5±0.8
−0.2±0.9
GPQA-D
42 / 43 / 44
51.5 / 50.7 / 50.0
51.0 / 50.4 / 51.2
−0.4 / −0.3 / +1.2
Mean ± SD
50.7±0.7
50.9±0.4
+0.2±0.9
HE+
42 / 43 / 44
87.0 / 87.0 / 86.1
88.1 / 85.3 / 86.1
+1.1 / −1.8 / +0.1
Mean ± SD
86.7±0.6
86.5±1.4
−0.2±1.4
Appendix
Table 24: OOD results for mixed-pool and source-filtered sampling. 30B Instruct → Qwen3-4B, M=48 , 25 optimization steps; selection seeds 42/43/44 share a fixed training seed. Scores are mean@8 (%); OOD average is the unweighted mean of GPQA-D, HE+, and LCB. Δ denotes paired Filtered − Mixed differences; means, SDs, and differences use unrounded scores.
Method
HE+
MBPP+
LCB
Code Macro
GPU ⋅ h
Uniform random
88.26 ± 0.32
71.33 ± 0.34
27.50 ± 0.43
62.36 ± 0.22
0
Semantic diversity
88.03
71.83
28.14
62.67
0.1
Hard selection
88.03
71.56
28.14
62.58
6.9
T–S disagreement
87.58
70.92
27.86
62.12
8.6
Cost-D-opt
88.42
71.17
27.64
62.41
48.6
Appendix
Table 25: Prompt selection within Eurus Code. Code-RL 4B → Qwen3-4B; M=48 , C=2,304 , 100 updates. Scores (%) use mean@8; Code Macro averages the three benchmarks. Uniform random reports mean ± sample SD across three support draws. Bold marks column-best scores. Costs cover selection only; 0 denotes no GPU screening.
Method
Math
GPQA-D
HE+
LCB
Avg 4
GPU ⋅ h
Uniform random
40.56 ± 1.07
49.87 ± 0.55
86.16 ± 0.52
25.74 ± 0.52
50.58 ± 0.61
0
Stratified random
41.79 ± 1.13
50.65 ± 1.03
86.81 ± 1.47
26.81 ± 0.36
51.52 ± 0.83
5.7
Semantic diversity
40.21
50.44
86.20
27.36
51.05
<0.1
Hard selection
40.00
49.87
87.65
26.64
51.04
5.7
Shortest
35.05
46.53
83.61
20.21
46.35
5.7
Appendix
Table 26: Prompt selection within DeepMath L6. 30B Instruct → Qwen3-4B; M=48 , C=2,304 , 25 updates. Scores (%) use mean@16 for Math and mean@8 otherwise; Avg 4 averages the four benchmarks. Random selectors report mean ± sample SD across three support draws. Bold marks column-best scores. Costs cover selection only; 0 denotes no GPU screening.
Method
Selection seed
E(S) (%)
Cosine
Avg4 (%)
Uniform random
42
29.7109
0.9670
37.5227
Uniform random
43
29.7557
0.9565
38.0760
Uniform random
44
23.4674
0.9722
37.3466
Stratified random
42
21.7985
0.9833
37.0561
Semantic diversity
42
23.5230
0.9721
37.4771
Hard selection
42
39.0314
0.9606
37.2021
Appendix
Table 27: Per-support gradient representativeness and transfer. JustRL–DeepMath, C=384 , M=8 . Each row is one support in Figure 8 , not a mean across selection seeds. E(S) and Avg4 are percentages; cosine is dimensionless.
Training cost
Performance (%)
Method
F/U
Generated
Loss
Time
Math
GPQA-D
HE+
LCB
Avg 4
DeepSeek 1.5B
Fresh-only
67/67
128.19
128.19
590.5
32.55
39.02
63.64
18.57
38.45
Fresh-only
20/20
39.13
39.13
178.8
31.09
38.83
63.80
18.36
38.02
Uniform token replay
17/68
32.78
57.64
205.1
32.92
39.65
62.73
19.21
38.63
Full-batch reuse
17/68
32.60
130.40
226.0
32.14
40.15
64.94
18.50
38.93
Appendix
Table 28: Performance and training cost with trajectory reuse. F/U denotes fresh rollout batches / optimizer updates; generated and loss tokens are in millions (including repeated exposures for loss), and time is in minutes. Math reports mean@16 and the other benchmarks mean@8; Avg 4 is their unweighted mean.