Black-box distillation is a practical route for transferring capabilities from API-accessible large language models that expose only text outputs into smaller student models. Recent on-policy adversarial methods such as GAD improve over SeqKD by forming an adversarial loop between a critic and a student, where the critic provides rewards for GRPO-based student policy optimization over the student's sampled responses. However, GRPO computes advantages from the within-group relative rewards of student samples for the same prompt, whereas the critic is trained primarily to distinguish teacher responses from student responses. This objective mismatch can produce reward groups with collapsed scale or fragile margins, leading to brittle grouped optimization signals. We propose Groupwise Reward Geometry Conditioning (GRGC), a two-stage framework that improves advantage construction by shaping student-side reward groups during both critic training and policy optimization. To improve critic-side conditioning, Gaussian groupwise Optimal Transport calibration regularizes the critic during training to produce reward groups with non-collapsed spread and smooth rank-wise gaps by matching sorted prompt-wise rewards to group-centered Gaussian quantiles. Building on this conditioned reward geometry, policy-side group power modulation reshapes the prompt-wise reward groups before they are converted into advantages, preserving the critic-induced ordering while increasing optimization-relevant margin separability. Extensive experiments across diverse teachers, student model families and scales, and training datasets demonstrate the effectiveness of GRGC on both in-distribution and out-of-distribution evaluations, while introducing negligible overhead over GAD. The code is available at https://github.com/2018cx/GRGC.
Figures & tables
Figure 1: Overview of GRGC. For each prompt x , the student samples a response group {G(x)i}i=1N , and the critic assigns raw sequence-level rewards {ri}i=1N . CGC regularizes prompt-group reward geometry during critic training, while PGM transforms the conditioned rewards into {r^i}i=1N before the final prompt-wise advantages {Ai}i=1N are computed for policy optimization.
Figure 2: Why BT-style adversarial critic training alone is insufficient for advantage construction. (a) Real GAD training often yields low-dispersion prompt groups. The x-axis denotes the group reward standard deviation normalized by the global reward standard deviation at the same step. (b)–(d) Toy grouped-bandit analysis. We fix latent utility scores ui to define the within-group preference order and set reward ri(a)=aui , so that varying a changes only the within-group reward dispersion. Smaller a means a more collapsed reward group. As a decreases, BT loss does not increase and may even decrease, while ranking becomes more noise-sensitive and grouped optimization deteriorates.
Rel. std <0.15
Rel. std <0.25
Stage
GAD
GRGC
GAD
GRGC
Early (500 steps)
8.57%
1.28%
15.73%
1.28%
Middle (500 steps)
8.69%
1.90%
16.04%
1.90%
Late (500 steps)
8.48%
1.86%
14.73%
1.86%
Table 1: Low-dispersion incidence over the complete 1500-step Qwen2.5-3B/GPT-5/LMSYS trajectories. Each stage contains 64000 groups per method.
Figure 3: Mechanism view of GRGC. Critic-side Gaussian OT calibration aligns prompt-group rewards with mean-centered Gaussian quantiles, promoting non-collapsed scale and smooth rank-wise gaps. Policy-side group power modulation remaps group rewards to enlarge informative margins while preserving the critic-induced ordering.
Model
Method
LMSYS
Dolly
SelfInst
Vicuna
Score
Win
Score
Win
Score
Win
Score
Win
GPT-5-Chat
Teacher
51.21
53.2%
49.57
49.4%
49.71
49.2%
50.27
55.0%
Qwen2.5-3B-Instruct
Before Distill.
45.79
11.9%
44.93
4.0%
46.56
12.8%
47.85
3.8%
SeqKD
47.06
18.4%
45.62
7.2%
46.83
16.9%
48.25
20.0%
GAD
48.44
25.1%
46.19
8.2%
47.33
16.9%
48.76
28.8%
GRGC
50.19
45.3%
47.40
26.6%
48.99
30.6%
50.48
38.8%
Table 2: Automatic evaluation results for models distilled from GPT-5-Chat and trained on the LMSYS-Chat training set. We report the averaged Qwen2.5-72B evaluation score and win rate.
Model
Method
LMSYS
Dolly
SelfInst
Vicuna
Score
Win
Score
Win
Score
Win
Score
Win
Doubao-Seed-2.0
Teacher
53.23
82.5%
53.14
75.8%
52.70
74.0%
53.20
93.8%
Qwen2.5-3B-Instruct
Before Distill.
45.79
11.9%
44.93
4.0%
46.56
12.8%
47.85
3.8%
SeqKD
46.95
25.3%
45.96
9.6%
46.87
16.9%
47.81
15.0%
GAD
48.16
42.2%
46.86
37.0%
47.78
38.0%
49.84
60.0%
GRGC
49.49
49.9%
47.92
48.0%
48.75
43.0%
51.23
68.8%
Table 3: Automatic evaluation results for models distilled from Doubao-Seed-2.0 and trained on the LMSYS-Chat training set. We report the averaged Qwen2.5-72B evaluation score and win rate.
Model
Method
LMSYS
Dolly
SelfInst
Vicuna
Score
Win
Score
Win
Score
Win
Score
Win
Doubao-Seed-2.0
Teacher
53.23
82.5%
53.14
75.8%
52.70
74.0%
53.20
93.8%
Qwen2.5-3B-Instruct
Before Distill.
45.79
11.9%
44.93
4.0%
46.56
12.8%
47.85
3.8%
SeqKD
46.65
31.9%
46.60
33.2%
46.63
28.9%
48.36
36.3%
GAD
47.92
33.6%
46.91
35.2%
48.38
36.4%
50.16
55.0%
GRGC
48.93
45.1%
47.88
42.8%
49.14
43.4%
51.25
62.5%
Table 4: Automatic evaluation results for models distilled from Doubao-Seed-2.0 and trained on Dolly Train. We report the averaged Qwen2.5-72B evaluation score and win rate on the test datasets.
Table 8Table 9Table 10
Figure 4: Mean relative prompt-group reward spread, computed as the mean group reward standard deviation normalized by the global reward standard deviation at each step. Our approach maintains a more stable reward scale.
Appendix figures & tables16 assets
Supplementary material from the paper’s appendix.
Appendix
Symbol
Definition
x
Input prompt.
y
Teacher response.
πT
Teacher policy.
πθ
Student policy.
θ
Student-policy parameters.
G(x)i
i -th student response.
Appendix
Table 11: Descriptions of symbols used in the main text and appendix.
Figure 5: The prompt wrapper for training and evaluation.
Figure 6: Automatic evaluation prompt.
λOT
0.001
0.005
0.01
0.02
0.05
0.1
Qwen2.5-3B-Instruct
48.92
50.10
50.19
50.22
50.16
49.94
Qwen2.5-1.5B-Instruct
46.79
47.27
47.36
47.31
47.32
47.17
Appendix
Table 12: Sensitivity to the critic-side OT weight λOT . Results are averaged Qwen2.5-72B evaluation scores. Very small λOT weakens geometry conditioning, while overly large λOT makes the critic too conservative. A broad middle range remains stable, and we use λOT=0.01 in the main experiments.
γ
1.0
1.2
1.4
1.5
1.6
2.0
Qwen2.5-3B-Instruct
47.52
50.20
50.15
50.19
50.14
49.72
Qwen2.5-1.5B-Instruct
46.40
47.29
47.34
47.36
47.30
47.02
Appendix
Table 13: Sensitivity to the power exponent γ . Results are averaged Qwen2.5-72B evaluation scores. Values too close to 1 reduce the shaping effect, whereas overly large values lead to overly sharp grouped rewards. A broad intermediate region is stable, and we use γ=1.5 in the main experiments.
Model
Method
LMSYS
Dolly
SelfInst
Vicuna
Score
Win
Score
Win
Score
Win
Score
Win
Doubao-Seed-2.0
Teacher
53.58
83.1%
53.47
82.4%
53.78
86.4%
53.63
93.8%
Qwen2.5-3B-Instruct
Before Distill.
45.79
11.9%
44.93
4.0%
46.56
12.8%
47.85
3.8%
SeqKD
46.09
41.8%
46.03
48.8%
46.81
46.3%
48.39
66.3%
GAD
47.85
51.4%
47.61
52.0%
48.24
51.7%
50.32
73.8%
GRGC
48.84
55.1%
48.98
61.4%
49.17
53.7%
51.85
82.5%
Appendix
Table 14: Automatic evaluation results for models distilled from Doubao-Seed-2.0 with an explicit step-by-step instruction and trained on the Dolly training set. We report the averaged Qwen2.5-72B evaluation score on the test datasets.
Figure 7: Human evaluation results on test sets. We compare GRGC to the models fine-tuned with SeqKD [ 22 ] and GAD [ 18 ] .
Figure 10: Unnormalized prompt-group reward spread . The solid lines show the mean prompt-group reward standard deviation, and shaded regions show the 10th–90th percentile range across prompt groups at each step. GRGC keeps the absolute reward spread in a more controlled range than GAD, confirming that the relative-spread improvement is not caused only by normalization with the global reward scale.
Method
Qwen2.5-3B-Instruct
Qwen2.5-1.5B-Instruct
Score
Win Rate
Score
Win Rate
GPT-5-Chat Teacher
66.13
76.8%
66.13
76.8%
Before Distill
54.75
49.5%
52.81
43.4%
SeqKD [ 22 ]
55.65
54.7%
54.20
46.8%
GAD [ 18 ]
56.41
50.1%
55.56
47.6%
Ours
60.52
60.8%
59.77
59.9%
Appendix
Table 15: GPT-OSS-120B evaluation results on the LMSYS test set. We report the averaged GPT-OSS-120B evaluation score and win rate.
Model
Teacher
Before Distill.
SeqKD
GAD
Ours
Qwen2.5-3B-Instruct
329.1
338.9
318.2
438.0
328.1
Qwen2.5-1.5B-Instruct
329.1
317.8
310.4
396.1
335.7
Appendix
Table 16: Average response length (token length) of different methods and models on the LMSYS test set. The models are trained using the LMSYS-Chat training set with the GPT-5-Chat teacher.
Student
Method
Strict
Loose
Qwen2.5-1.5B
SeqKD
54.56
57.79
GAD
54.32
58.51
GRGC
56.12
60.79
Qwen2.5-3B
SeqKD
66.91
70.14
GAD
68.35
72.06
GRGC
70.98
75.54
Appendix
Table 17: Rule-based IFEval accuracy (%) under strict and loose verification.
Method
Minerva Math
Gaokao-MathQA
Before Distillation
11.40
43.53
SeqKD
13.97
43.82
GAD
15.07
45.59
GRGC
17.65
49.12
Appendix
Table 18: Accuracy (%) of Qwen2.5-1.5B after long-form mathematical-reasoning distillation on 30000 OpenR1-Math-220k problems.
Method
Before Distill
SeqKD
GAD
Ours
Accuracy
46.4%
39.8%
44.0%
49.6%
Appendix
Table 19: Math500 evaluation for Qwen2.5-1.5B-Instruct distilled from GPT-5-Chat on LMSYS-Chat. Only 3.4% of LMSYS-Chat training prompts are math-related, making this an out-of-domain reasoning stress test rather than a math-specialized training setting.
Method
Seed=0
Seed=1
Seed=2
Seed=3
Seed=4
Mean
Std
95% CI
SeqKD
47.06
47.10
47.05
47.04
47.07
47.06
0.02
[47.04, 47.08]
GAD
48.44
48.44
48.43
48.38
48.43
48.42
0.03
[48.38, 48.46]
Ours
50.19
50.20
50.15
50.22
50.18
50.19
0.03
[50.15, 50.23]
Appendix
Table 20: Multi-seed Qwen2.5-72B automatic evaluation score. Seed=0 is the default evaluation seed used in the main results. The 95% Confidence Interval (CI) represents the range within which the true mean score is expected to lie with 95% confidence.
Method
Seed=0
Seed=1
Seed=2
Seed=3
Seed=4
Mean
Std
95% CI
SeqKD
18.4
18.6
18.4
18.8
18.2
18.46
0.24
[18.16, 18.75]
GAD
25.1
25.5
24.6
24.8
24.8
24.97
0.32
[24.58, 25.36]
Ours
45.3
45.3
44.9
45.3
46.3
45.43
0.54
[44.75, 46.10]
Appendix
Table 21: Multi-seed Qwen2.5-72B automatic evaluation win rate. Seed=0 is the default evaluation seed used in the main results. The 95% Confidence Interval (CI) represents the range within which the true mean score is expected to lie with 95% confidence.
Method
Time complexity
Memory complexity
Dominant terms
GAD
O(BFπroll(L))+O(BFD(L))+O(BFπupd(L))+O(B)
Mπ+MD+O(B)
rollout critic forward/backward actor update
GRGC (ours)
O(BFπroll(L))+O(BFD(L))+O(BFπupd(L))+O(BlogN)
Mπ+MD+O(B)
rollout + critic + actor update OT sorting/matching group power shaping
Appendix
Table 22: Asymptotic training complexity comparison. Here B is the number of sampled student responses in a minibatch, N is the prompt-group size, and L is the average response length. Since N is a small fixed group size in our setting, grouped operations with complexity O(BlogN) behave as low-order overheads and do not change the dominant model-scale training complexity. As a result, the total wall-clock training time of GRGC is expected to remain nearly identical to that of GAD.
University of Chinese Academy of Sciences · 2Chinese Information Processing Laboratory Institute of Software, Chinese Academy of Sciences 2University of Chinese Academy of Sciences, Beijing, China · 3Xiaohongshu