cs.LGOct 8, 2026

PIVOT: Perplexity-Informed KD-to-RL Transition Scheduling for Vertical-Domain Few-Shot Distillation

Authors: Heng Li, Yong Zhang, Ning Cheng, Zhigen Li, Yun Zhu, Yanmeng Wang, Shaojun Wang, Jing Xiao

Organizations: Ping An Technology (Shenzhen) Co., Ltd., China · University of Science and Technology of China

Abstract

Vertical-domain few-shot classification remains challenging for small language models, as limited supervision makes it difficult to acquire domain-specific decision knowledge. On-Policy Distillation (OPD) can improve teacher-guided adaptation by supervising student-generated rollouts, while GRPO-based reinforcement learning can further refine downstream predictions. However, existing KD-to-RL pipelines typically rely on globally fixed transition schedules, ignoring that different samples may require different amounts of teacher-guided acquisition before reward-driven refinement. We propose PIVOT (Perplexity-Informed Transition Optimization), a dynamic transition framework that routes samples between OPD and GRPO according to teacher-evaluated sequence perplexity. PIVOT moves low-perplexity samples to GRPO for reward-driven refinement while keeping high-perplexity samples under OPD for continued domain knowledge acquisition. Experiments on Banking77 and HWU64 show that PIVOT consistently outperforms continued OPD and globally synchronized OPD→\rightarrowGRPO baselines under the same number of post-warm-up student optimization steps, achieving stronger downstream performance and more stable training dynamics.

Figures & tables

Appendix figures & tables2 assets

Supplementary material from the paper’s appendix.

Appendix

Explore similar work

CardsList
  1. PIVOT: Pivot-Aware On Policy Self Distillation for Multi-Turn VLM Agents

    Sep 28, 2026Jiazhou Zhou, Hu Zhou, Yucheng Chen +3On-Policy Self-DistillationReinforcement Learning with Verifiable Rewards

  2. Eliciting Weak-to-Strong Generalization with On-Policy Reverse Distillation

    Sep 8, 2026Youngrok Park, Sangmin Bae, Hojung Jung +6Policy GradientWeak-to-Strong Generalization