Vertical-domain few-shot classification remains challenging for small language models, as limited supervision makes it difficult to acquire domain-specific decision knowledge. On-Policy Distillation (OPD) can improve teacher-guided adaptation by supervising student-generated rollouts, while GRPO-based reinforcement learning can further refine downstream predictions. However, existing KD-to-RL pipelines typically rely on globally fixed transition schedules, ignoring that different samples may require different amounts of teacher-guided acquisition before reward-driven refinement. We propose PIVOT (Perplexity-Informed Transition Optimization), a dynamic transition framework that routes samples between OPD and GRPO according to teacher-evaluated sequence perplexity. PIVOT moves low-perplexity samples to GRPO for reward-driven refinement while keeping high-perplexity samples under OPD for continued domain knowledge acquisition. Experiments on Banking77 and HWU64 show that PIVOT consistently outperforms continued OPD and globally synchronized OPD→GRPO baselines under the same number of post-warm-up student optimization steps, achieving stronger downstream performance and more stable training dynamics.
Figures & tables
Figure 1: Overview of PIVOT. PIVOT estimates sample-level acquisition states using teacher-side perplexity over on-policy student rollouts, and progressively performs asynchronous KD-to-RL transition by routing low-perplexity samples to GRPO refinement while retaining high-perplexity samples under continued OPD acquisition.
Figure 2: Training dynamics on Banking77 under aligned total optimization steps. For globally synchronized OPD → GRPO, the x-axis denotes cumulative optimization steps across the OPD and GRPO stages, and the purple star marks the global transition point. Compared with continued OPD and globally synchronized OPD → GRPO, our progressive sample-level transition achieves higher final accuracy with substantially smoother reward, entropy, and completion-length dynamics.
Method
Banking77
HWU64
Teacher
73.25
79.65
SeqKD
58.70
70.17
GRPO
67.66
77.60
OPD
72.47
79.55
OPD → GRPO
78.05
83.09
PIVOT
80.26
85.04
Table 1: Main results under 480 post-warm-up student optimization steps. All results report test-set classification accuracy.
Dataset
Setting
Student
Teacher
OPD → GRPO
PIVOT
Δ
Banking77
Default
0.5B
Qwen2.5-7B
78.05
80.26
+2.21
Banking77
Larger student
1.5B
Qwen2.5-7B
80.00
82.82
+2.82
Banking77
Stronger teacher
0.5B
DianJin-32B
80.30
81.33
+1.03
HWU64
Default
0.5B
Qwen2.5-7B
83.09
85.04
+1.95
HWU64
Larger student
1.5B
Qwen2.5-7B
82.53
85.97
+3.44
HWU64
Stronger teacher
0.5B
Qwen3-32B
84.39
86.71
+2.32
Table 2: Scaling analysis across different student and teacher model configurations. Δ denotes the improvement of PIVOT over OPD → GRPO.
Dataset
Low-PPL
Random
High-PPL
Banking77
80.26
78.31
77.27
HWU64
85.04
84.11
83.46
Table 3: Routing-rule ablation under progressive OPD-to-GRPO transition. Column names indicate which samples are preferentially routed to GRPO earlier during training.
Transition Strategy
ACC
Global OPD → GRPO
78.05
Linear progressive routing
79.58
Quadratic progressive routing
79.19
Cosine progressive routing
80.26
Table 4: Transition-schedule ablation on Banking77 under the same number of student optimization steps.
Figure 3: Pass@32 performance grouped by teacher-side perplexity buckets during OPD training. High-perplexity samples exhibit substantially larger relative improvement throughout training, whereas low-perplexity samples saturate much earlier, suggesting that the optimal transition timing from OPD to RL varies across samples.
Appendix figures & tables2 assets
Supplementary material from the paper’s appendix.
Appendix
Method
Banking77
SFT-ZeroShot
60.30
DataAug-SFT
71.75
Continued OPD
72.47
OPD → GRPO
78.05
PIVOT
80.26
Appendix
Table 7: Comparison between teacher-generated data augmentation and different OPD/GRPO training strategies on Banking77. DataAug-SFT is trained on 5,119 teacher-filtered synthetic examples generated from the same 5-shot split.
Hyperparameter
Value
Total optimization steps
480
Batch size
256
Rollouts / group size
8
Warmup ratio
0.05
Max sequence length
4096
Max new tokens
512
Appendix
Table 9: Training hyperparameters used in all experiments.
AI Thrust, The Hong Kong University of Science and Technology (Guangzhou) · International Digital Economy Academy (IDEA) · The Hong Kong Polytechnic University +2