Black-box distillation is a practical route for transferring capabilities from API-accessible large language models that expose only text outputs into smaller student models. Recent on-policy adversarial methods such as GAD improve over SeqKD by forming an adversarial loop between a critic and a student, where the critic provides rewards for GRPO-based student policy optimization over the student's sampled responses. However, GRPO computes advantages from the within-group relative rewards of student samples for the same prompt, whereas the critic is trained primarily to distinguish teacher responses from student responses. This objective mismatch can produce reward groups with collapsed scale or fragile margins, leading to brittle grouped optimization signals. We propose Groupwise Reward Geometry Conditioning (GRGC), a two-stage framework that improves advantage construction by shaping student-side reward groups during both critic training and policy optimization. To improve critic-side conditioning, Gaussian groupwise Optimal Transport calibration regularizes the critic during training to produce reward groups with non-collapsed spread and smooth rank-wise gaps by matching sorted prompt-wise rewards to group-centered Gaussian quantiles. Building on this conditioned reward geometry, policy-side group power modulation reshapes the prompt-wise reward groups before they are converted into advantages, preserving the critic-induced ordering while increasing optimization-relevant margin separability. Extensive experiments across diverse teachers, student model families and scales, and training datasets demonstrate the effectiveness of GRGC on both in-distribution and out-of-distribution evaluations, while introducing negligible overhead over GAD. The code is available at https://github.com/2018cx/GRGC.
Figures & tables
Figure 1: Overview of GRGC. For each prompt x , the student samples a response group {G(x)i}i=1N , and the critic assigns raw sequence-level rewards {ri}i=1N . CGC regularizes prompt-group reward geometry during critic training, while PGM transforms the conditioned rewards into {r^i}i=1N before the final prompt-wise advantages {Ai}i=1N are computed for policy optimization.
Figure 2: Why BT-style adversarial critic training alone is insufficient for advantage construction. (a) Real GAD training often yields low-dispersion prompt groups. The x-axis denotes the group reward standard deviation normalized by the global reward standard deviation at the same step. (b)–(d) Toy grouped-bandit analysis. We fix latent utility scores ui to define the within-group preference order and set reward ri(a)=aui , so that varying a changes only the within-group reward dispersion. Smaller a means a more collapsed reward group. As a decreases, BT loss does not increase and may even decrease, while ranking becomes more noise-sensitive and grouped optimization deteriorates.
Rel. std <0.15
Rel. std <0.25
Stage
GAD
GRGC
GAD
GRGC
Early (500 steps)
8.57%
1.28%
15.73%
1.28%
Middle (500 steps)
8.69%
1.90%
16.04%
1.90%
Late (500 steps)
8.48%
1.86%
14.73%
1.86%
Table 1: Low-dispersion incidence over the complete 1500-step Qwen2.5-3B/GPT-5/LMSYS trajectories. Each stage contains 64000 groups per method.
Figure 3: Mechanism view of GRGC. Critic-side Gaussian OT calibration aligns prompt-group rewards with mean-centered Gaussian quantiles, promoting non-collapsed scale and smooth rank-wise gaps. Policy-side group power modulation remaps group rewards to enlarge informative margins while preserving the critic-induced ordering.
Model
Method
LMSYS
Dolly
SelfInst
Vicuna
Score
Win
Score
Win
Score
Win
Score
Win
GPT-5-Chat
Teacher
51.21
53.2%
49.57
49.4%
49.71
49.2%
50.27
55.0%
Qwen2.5-3B-Instruct
Before Distill.
45.79
11.9%
44.93
4.0%
46.56
12.8%
47.85
3.8%
SeqKD
47.06
18.4%
45.62
7.2%
46.83
16.9%
48.25
20.0%
GAD
48.44
25.1%
46.19
8.2%
47.33
16.9%
48.76
28.8%
GRGC
50.19
45.3%
47.40
26.6%
48.99
30.6%
50.48
38.8%
Table 2: Automatic evaluation results for models distilled from GPT-5-Chat and trained on the LMSYS-Chat training set. We report the averaged Qwen2.5-72B evaluation score and win rate.
Model
Method
LMSYS
Dolly
SelfInst
Vicuna
Score
Win
Score
Win
Score
Win
Score
Win
Doubao-Seed-2.0
Teacher
53.23
82.5%
53.14
75.8%
52.70
74.0%
53.20
93.8%
Qwen2.5-3B-Instruct
Before Distill.
45.79
11.9%
44.93
4.0%
46.56
12.8%
47.85
3.8%
SeqKD
46.95
25.3%
45.96
9.6%
46.87
16.9%
47.81
15.0%
GAD
48.16
42.2%
46.86
37.0%
47.78
38.0%
49.84
60.0%
GRGC
49.49
49.9%
47.92
48.0%
48.75
43.0%
51.23
68.8%
Table 3: Automatic evaluation results for models distilled from Doubao-Seed-2.0 and trained on the LMSYS-Chat training set. We report the averaged Qwen2.5-72B evaluation score and win rate.
Model
Method
LMSYS
Dolly
SelfInst
Vicuna
Score
Win
Score
Win
Score
Win
Score
Win
Doubao-Seed-2.0
Teacher
53.23
82.5%
53.14
75.8%
52.70
74.0%
53.20
93.8%
Qwen2.5-3B-Instruct
Before Distill.
45.79
11.9%
44.93
4.0%
46.56
12.8%
47.85
3.8%
SeqKD
46.65
31.9%
46.60
33.2%
46.63
28.9%
48.36
36.3%
GAD
47.92
33.6%
46.91
35.2%
48.38
36.4%
50.16
55.0%
GRGC
48.93
45.1%
47.88
42.8%
49.14
43.4%
51.25
62.5%
Table 4: Automatic evaluation results for models distilled from Doubao-Seed-2.0 and trained on Dolly Train. We report the averaged Qwen2.5-72B evaluation score and win rate on the test datasets.
Table 8Table 9Table 10
Figure 4: Mean relative prompt-group reward spread, computed as the mean group reward standard deviation normalized by the global reward standard deviation at each step. Our approach maintains a more stable reward scale.
Appendix figures & tables16 assets
Supplementary material from the paper’s appendix.
Appendix
Symbol
Definition
x
Input prompt.
y
Teacher response.
πT
Teacher policy.
πθ
Student policy.
θ
Student-policy parameters.
G(x)i
i -th student response.
Appendix
Table 11: Descriptions of symbols used in the main text and appendix.
Figure 5: The prompt wrapper for training and evaluation.
Figure 6: Automatic evaluation prompt.
λOT
0.001
0.005
0.01
0.02
0.05
0.1
Qwen2.5-3B-Instruct
48.92
50.10
50.19
50.22
50.16
49.94
Qwen2.5-1.5B-Instruct
46.79
47.27
47.36
47.31
47.32
47.17
Appendix
Table 12: Sensitivity to the critic-side OT weight λOT . Results are averaged Qwen2.5-72B evaluation scores. Very small λOT weakens geometry conditioning, while overly large λOT makes the critic too conservative. A broad middle range remains stable, and we use λOT=0.01 in the main experiments.
γ
1.0
1.2
1.4
1.5
1.6
2.0
Qwen2.5-3B-Instruct
47.52
50.20
50.15
50.19
50.14
49.72
Qwen2.5-1.5B-Instruct
46.40
47.29
47.34
47.36
47.30
47.02
Appendix
Table 13: Sensitivity to the power exponent γ . Results are averaged Qwen2.5-72B evaluation scores. Values too close to 1 reduce the shaping effect, whereas overly large values lead to overly sharp grouped rewards. A broad intermediate region is stable, and we use γ=1.5 in the main experiments.
Model
Method
LMSYS
Dolly
SelfInst
Vicuna
Score
Win
Score
Win
Score
Win
Score
Win
Doubao-Seed-2.0
Teacher
53.58
83.1%
53.47
82.4%
53.78
86.4%
53.63
93.8%
Qwen2.5-3B-Instruct
Before Distill.
45.79
11.9%
44.93
4.0%
46.56
12.8%
47.85
3.8%
SeqKD
46.09
41.8%
46.03
48.8%
46.81
46.3%
48.39
66.3%
GAD
47.85
51.4%
47.61
52.0%
48.24
51.7%
50.32
73.8%
GRGC
48.84
55.1%
48.98
61.4%
49.17
53.7%
51.85
82.5%
Appendix
Table 14: Automatic evaluation results for models distilled from Doubao-Seed-2.0 with an explicit step-by-step instruction and trained on the Dolly training set. We report the averaged Qwen2.5-72B evaluation score on the test datasets.
Figure 7: Human evaluation results on test sets. We compare GRGC to the models fine-tuned with SeqKD [ 22 ] and GAD [ 18 ] .
Figure 10: Unnormalized prompt-group reward spread . The solid lines show the mean prompt-group reward standard deviation, and shaded regions show the 10th–90th percentile range across prompt groups at each step. GRGC keeps the absolute reward spread in a more controlled range than GAD, confirming that the relative-spread improvement is not caused only by normalization with the global reward scale.
Method
Qwen2.5-3B-Instruct
Qwen2.5-1.5B-Instruct
Score
Win Rate
Score
Win Rate
GPT-5-Chat Teacher
66.13
76.8%
66.13
76.8%
Before Distill
54.75
49.5%
52.81
43.4%
SeqKD [ 22 ]
55.65
54.7%
54.20
46.8%
GAD [ 18 ]
56.41
50.1%
55.56
47.6%
Ours
60.52
60.8%
59.77
59.9%
Appendix
Table 15: GPT-OSS-120B evaluation results on the LMSYS test set. We report the averaged GPT-OSS-120B evaluation score and win rate.
Model
Teacher
Before Distill.
SeqKD
GAD
Ours
Qwen2.5-3B-Instruct
329.1
338.9
318.2
438.0
328.1
Qwen2.5-1.5B-Instruct
329.1
317.8
310.4
396.1
335.7
Appendix
Table 16: Average response length (token length) of different methods and models on the LMSYS test set. The models are trained using the LMSYS-Chat training set with the GPT-5-Chat teacher.
Student
Method
Strict
Loose
Qwen2.5-1.5B
SeqKD
54.56
57.79
GAD
54.32
58.51
GRGC
56.12
60.79
Qwen2.5-3B
SeqKD
66.91
70.14
GAD
68.35
72.06
GRGC
70.98
75.54
Appendix
Table 17: Rule-based IFEval accuracy (%) under strict and loose verification.
Method
Minerva Math
Gaokao-MathQA
Before Distillation
11.40
43.53
SeqKD
13.97
43.82
GAD
15.07
45.59
GRGC
17.65
49.12
Appendix
Table 18: Accuracy (%) of Qwen2.5-1.5B after long-form mathematical-reasoning distillation on 30000 OpenR1-Math-220k problems.
Method
Before Distill
SeqKD
GAD
Ours
Accuracy
46.4%
39.8%
44.0%
49.6%
Appendix
Table 19: Math500 evaluation for Qwen2.5-1.5B-Instruct distilled from GPT-5-Chat on LMSYS-Chat. Only 3.4% of LMSYS-Chat training prompts are math-related, making this an out-of-domain reasoning stress test rather than a math-specialized training setting.
Method
Seed=0
Seed=1
Seed=2
Seed=3
Seed=4
Mean
Std
95% CI
SeqKD
47.06
47.10
47.05
47.04
47.07
47.06
0.02
[47.04, 47.08]
GAD
48.44
48.44
48.43
48.38
48.43
48.42
0.03
[48.38, 48.46]
Ours
50.19
50.20
50.15
50.22
50.18
50.19
0.03
[50.15, 50.23]
Appendix
Table 20: Multi-seed Qwen2.5-72B automatic evaluation score. Seed=0 is the default evaluation seed used in the main results. The 95% Confidence Interval (CI) represents the range within which the true mean score is expected to lie with 95% confidence.
Method
Seed=0
Seed=1
Seed=2
Seed=3
Seed=4
Mean
Std
95% CI
SeqKD
18.4
18.6
18.4
18.8
18.2
18.46
0.24
[18.16, 18.75]
GAD
25.1
25.5
24.6
24.8
24.8
24.97
0.32
[24.58, 25.36]
Ours
45.3
45.3
44.9
45.3
46.3
45.43
0.54
[44.75, 46.10]
Appendix
Table 21: Multi-seed Qwen2.5-72B automatic evaluation win rate. Seed=0 is the default evaluation seed used in the main results. The 95% Confidence Interval (CI) represents the range within which the true mean score is expected to lie with 95% confidence.
Method
Time complexity
Memory complexity
Dominant terms
GAD
O(BFπroll(L))+O(BFD(L))+O(BFπupd(L))+O(B)
Mπ+MD+O(B)
rollout critic forward/backward actor update
GRGC (ours)
O(BFπroll(L))+O(BFD(L))+O(BFπupd(L))+O(BlogN)
Mπ+MD+O(B)
rollout + critic + actor update OT sorting/matching group power shaping
Appendix
Table 22: Asymptotic training complexity comparison. Here B is the number of sampled student responses in a minibatch, N is the prompt-group size, and L is the average response length. Since N is a small fixed group size in our setting, grouped operations with complexity O(BlogN) behave as low-order overheads and do not change the dominant model-scale training complexity. As a result, the total wall-clock training time of GRGC is expected to remain nearly identical to that of GAD.
Reinforcement learning with verifiable rewards (RLVR) has become a standard paradigm for post-training large language models (LLMs). While Group Relative Policy Optimization (GRPO) is widely adopted, it suffers from sparse reward signals and loses gradients entirely when all responses within a group receive identical rewards. On-policy distillation (OPD) offers a natural remedy by providing dense, token-level supervision from a teacher model. However, naively combining GRPO with OPD leads to degraded performance, due to three underlying causes: not all samples benefit from distillation; fitting too quickly to the teacher undermines the exploratory capacity of RL; and OPD's advantages are asymmetric, suppressing most tokens. To address these challenges, we propose RSTG (Recovering Learning Signals via Adaptive Teacher Guidance), which applies distillation selectively and precisely where it matters most. At the sample level, OPD is restricted to negative zero-variance prompts with each sample weighted by the teacher's confidence score. At the token level, distillation targets only tokens with high student entropy or large teacher-student divergence. We further augment training with SFT on correct trajectories generated by the teacher model, injecting positive gradient signals where RL yields none. Experiments demonstrate that RSTG substantially outperforms naive GRPO+OPD by +4.02% on math and +3.05% on code.
Zhuowen Han, Jinwei Xiao, Zhengxi Lu +9
TJUNLP Lab, School of Computer Science and Technology, Tianjin University · Meituan Longcat Team
On-policy distillation transfers reasoning capabilities by training a student model on its own generated trajectories using token-level feedback from a teacher. However, we identify a critical bottleneck, \textbf{Supervision Fidelity Decay (SFD)}: as student-generated prefixes lengthen, the teacher's next-token distribution becomes less confident and less discriminative. Consequently, the teacher-dependent corrective signal in reverse-KL distillation weakens, causing student drift to compound across long reasoning chains. To mitigate SFD, we introduce \textbf{Lookahead Group Reward (\ours{})}. Building on the insight that next-step teacher confidence reflects the discriminative strength of future reverse-KL supervision, \ours{} evaluates the student's top-K candidate tokens by the teacher confidence they induce at the subsequent step and assigns a group-normalized reward. To maintain computational efficiency, we further design an entropy-triggered tree-attention mechanism. Across six math and code benchmarks, \ours{} improves mean@8 by \textbf{2.57} points over OPD for a 7B student, with gains increasing in longer-generation and reaching +\textbf{4.92} points on AIME-26 at 39k tokens.
Yanjiang Liu, Jie Lou, Xinyan Guan +7
University of Chinese Academy of Sciences · 2Chinese Information Processing Laboratory Institute of Software, Chinese Academy of Sciences 2University of Chinese Academy of Sciences, Beijing, China · 3Xiaohongshu
On-policy distillation is a powerful way to transfer reasoning ability from a strong teacher to a smaller student: the student samples trajectories from its own policy, and the teacher provides dense token-level supervision on the states the student actually visits. However, this supervision is not always reliable: a teacher can assign high likelihood to plausible but incorrect solutions, or low likelihood to correct student solutions that follow different reasoning paths. Unconditionally distilling the teacher can therefore reinforce bad modes or erase useful student behavior. To address these limitations, we introduce RG-OPD: Reward-Gated On-Policy Distillation that uses verifier feedback to decide when teacher logits should be trusted. RG-OPD bridges sparse verifier rewards and dense teacher logits, preserving token-level supervision while filtering misleading teacher signals. Across reasoning and coding benchmarks, RG-OPD produces stronger distilled students, outperforming both vanilla reverse-KL distillation and the recent TSD-KD baseline. At 1K generation length, RG-OPD improves over reverse-KL by 2.9 points and over TSD-KD by 4.9 points; in the long-generation setting, it improves over the untuned student by 8.2 points. Our code is available at https://github.com/UoC-tail/RG-OPD.
Mohammad Sadegh Akhondzadeh, Vijay Lingam, Atula Tejaswi +3