Verbal feedback can identify errors and prescribe corrections, providing rich supervision for language-model post-training even when reliable programmatic verifiers are unavailable. Such feedback, often generated by a capable model, can be used to condition the teacher in on-policy distillation, which trains the student to match the teacher's predictions on student-generated rollouts. However, this approach can transfer teacher preferences that the feedback did not motivate, while leaving much of the feedback's guidance unused. We find that both problems come from the standard on-policy distillation objective, specifically the divergence it minimizes and the distribution it uses as its target. Our proposed method addresses both limitations. First, to isolate the information conveyed by the feedback from the teacher's inherent preferences, we treat verbal feedback as evidence for or against the hypothesis that a particular token comes next at a given prefix. We then adopt a probabilistic confirmation framework which uniquely determines an ordering over the vocabulary based on the teacher's predictions before and after it receives feedback. Using a confirmation score consistent with this ordering, we construct a target distribution within a trust region of the student. Second, to learn from guidance that student rollouts can leave unused, we derive a simple shared-rollout estimator of a symmetric divergence between the student and target distributions over rollouts, reusing student and feedback-conditioned teacher rollouts in both directions through importance weighting. Empirical evaluations show that our method outperforms the common on-policy distillation recipe and a recent contrastive variant on knowledge-based and agentic benchmarks.
Figures & tables
Method
Trivia Fantasy
EmbodiedEval
SciKnowEval (L3)
ToolAlpaca
Chem.
Phys.
Biology
Materials
Base model
02.8±1.9
57.3±2.0
31.9±0.3
53.8±0.1
22.1±4.1
68.4±0.8
47.6±1.7
SDPO
61.0±7.1
73.7±1.0
35.9±3.4
54.9±1.9
27.5±0.7
73.2±0.3
59.5±0.5
W2S-OPD
41.7±7.2
60.8±4.2
35.0±1.1
54.8±0.4
27.3±1.9
72.1±0.4
59.7±2.0
Ours
98.3±2.9
82.1±3.0
50.2±2.9
66.9±1.6
28.2±1.8
73.7±1.5
61.9±2.4
Table 1: Our method achieves the highest mean score across all benchmarks and SciKnowEval subjects. Scores (%; higher is better) are means with 95% confidence intervals; the highest mean in each column is bold. Trained-model scores are averaged over three checkpoint evaluations; base-model scores are averaged over independent evaluation passes (Section 5.1 ).
Appendix figures & tables3 assets
Supplementary material from the paper’s appendix.
Appendix
Policy pair
Vocabulary size, horizon
Relative gap
Neural, KL calibration 0.02 nats
32,768,512
2.0%
Neural, KL calibration 0.5 nats
32,768,512
3.3%
Appendix
Table 2: Relative reduction from unskewed Jeffreys for perturbed neural policies. Entries are Monte Carlo estimates using 4,096 rollouts per policy, shared prefixes for both objectives, and exact vocabulary sums.
Trivia Fantasy
EmbodiedEval
SciKnowEval (L3)
ToolAlpaca
Student
Qwen3.5-2B
Qwen3.5-4B
Qwen3.5-4B
Qwen3.5-4B
Training set
20 questions
3,765 scenarios
450–1,890 per subject
4,046 requests
Evaluation set
20 questions
12 suites × 20 tasks
50 / 210 / 80 / 94 (bio. / chem. / phys. / mat.)
68 requests
Learning rate
10−6
10−6
10−7
5×10−7
Warmup steps
5
50
50
50
Sequences per update
512
128
128
128
Appendix
Table 3: Settings shared by all methods on each benchmark. All methods use Adam with weight decay 0.01 , gradient-norm clipping at 1 , a constant learning rate after linear warmup, training sampling at temperature 1 and top- p0.95 , and Gemini 3.7 Flash at temperature 1 as the feedback judge. A rollout becomes stale once the policy that generated it is more than four weight synchronizations old.
Ours
SDPO
W2S-OPD
Teacher
frozen initial student
frozen initial student
frozen initial student
Teacher contexts
with and without feedback
with feedback
with and without feedback
Training rollouts
student and teacher, λ=1/2
student only
student only
Token divergence
α -skew Jeffreys, α=0.01
reverse KL
reverse KL
Target
corrected target equation 6 , β=10
feedback-conditioned teacher
pW2S , contrast strength 0.75 , anchor π0
Importance weights
prefix weights equation 10 , forward weight clipped at cIS=2
Large language models are often post-trained with sparse verifier rewards, which indicate whether a sampled trajectory succeeds but provide limited guidance about where reasoning succeeds or fails. On-policy distillation (OPD) offers denser token-level supervision by training on student-generated trajectories, yet existing methods typically distill each rollout independently and ignore the other attempts sampled for the same prompt. We introduce Multi-Rollout On-Policy Distillation (MOPD), a peer-conditioned distillation framework that uses the student's local rollout group to construct more informative teacher signals. MOPD conditions the teacher on both successful and failed peer rollouts: successes provide positive evidence for valid reasoning patterns, while failures provide structured negative evidence about plausible mistakes to avoid. We study two peer-context constructions: positive peer imitation and contrastive success-failure conditioning. Experiments on competitive programming, mathematical reasoning, scientific question answering, and tool-use benchmarks show that MOPD consistently improves over standard on-policy baselines. Further teacher-signal analysis shows that mixed success-failure contexts better align teacher scores with verifier rewards, indicating that the gains arise from more faithful, instance-adaptive supervision. These results indicate that effective on-policy distillation should exploit the student's multi-rollout trial-and-error behavior rather than treating rollouts as isolated samples.
Weichen Yu, Xiaomin Li, Yizhou Zhao +8
1Carnegie Mellon University · 2Microsoft · 3Purdue University
Reinforcement learning from verifiable rewards (RLVR) suffers from sparse outcome signals, creating severe exploration bottlenecks on complex reasoning tasks. Recent on-policy self-distillation methods attempt to address this by utilizing language feedback to generate dense, token-level supervision. However, these approaches rely on a fixed, passive teacher to interpret the feedback. As the student policy improves, the teacher's zero-shot assessment capabilities plateau, ultimately halting further learning. To overcome this, we propose Variational Policy Distillation (VPD), a framework that formalizes learning from language feedback as a Variational Expectation-Maximization (EM) problem. VPD co-evolves both policies: in the E-step, the teacher is actively refined on trajectory outcomes via an adaptive trust-region update, translating textual feedback into a dynamically improved target token distribution. In the M-step, the student internalizes this dense distributional guidance on its own on-policy rollouts. By continuously improving the teacher's ability to extract actionable signals from textual critique, VPD overcomes the limitations of passive distillation. Evaluated across diverse sources of diagnostic feedback on scientific reasoning and code generation tasks, VPD consistently outperforms both standard RLVR and existing self-distillation baselines. Finally, by stress-testing our framework on rigid mathematical reasoning and cold-start regimes, we illuminate the fundamental bounds of feedback-driven self-distillation compared to pure environment-driven RL.
Knowledge distillation transfers reasoning capabilities from large teachers to efficient students. However, token-level on-policy distillation (OPD) constrains student exploration and requires teacher token probabilities, precluding distillation from black-box teachers that provide only text outputs. We introduce On-policy Verbal Distillation (OVD), a framework that uses verbal scores from black-box teachers to rank student-generated sub-trajectories, retaining high-scoring ones and replacing low-scoring ones with teacher-generated continuations. We analyze when ranking induced by verbal scores can guide distribution approximation: under a density-ratio calibration condition on acceptance probabilities and bounded teacher-replacement error, we bound the approximation error between the resulting mixed trajectory distribution and a teacher-preferred target. On Web Q&A, OVD achieves 41.09% average EM with teacher feedback at inference, exceeding the strongest evaluated baseline by 5.89 percentage points. On AMC23, OVD-FR improves accuracy over RLVR by 10.0 percentage points (52.5% to 62.5%) after 600 training steps on 128 problems. Further experiments suggest that retaining student-generated prefixes helps preserve exploration and mitigate trajectory-level entropy collapse. OVD also improves training efficiency: resampling selected suffixes rather than entire responses reduces mean per-step training time by 10.2% in the 128-problem setting. Project page: https://menik1126.github.io/ovd-project-page/.
Jing Xiong, Hui Shen, Shansan Gong +7
The University of Hong Kong, Hong Kong, China · Nanjing University, Nanjing, China · Huawei Technologies, China