Verbal feedback can identify errors and prescribe corrections, providing rich supervision for language-model post-training even when reliable programmatic verifiers are unavailable. Such feedback, often generated by a capable model, can be used to condition the teacher in on-policy distillation, which trains the student to match the teacher's predictions on student-generated rollouts. However, this approach can transfer teacher preferences that the feedback did not motivate, while leaving much of the feedback's guidance unused. We find that both problems come from the standard on-policy distillation objective, specifically the divergence it minimizes and the distribution it uses as its target. Our proposed method addresses both limitations. First, to isolate the information conveyed by the feedback from the teacher's inherent preferences, we treat verbal feedback as evidence for or against the hypothesis that a particular token comes next at a given prefix. We then adopt a probabilistic confirmation framework which uniquely determines an ordering over the vocabulary based on the teacher's predictions before and after it receives feedback. Using a confirmation score consistent with this ordering, we construct a target distribution within a trust region of the student. Second, to learn from guidance that student rollouts can leave unused, we derive a simple shared-rollout estimator of a symmetric divergence between the student and target distributions over rollouts, reusing student and feedback-conditioned teacher rollouts in both directions through importance weighting. Empirical evaluations show that our method outperforms the common on-policy distillation recipe and a recent contrastive variant on knowledge-based and agentic benchmarks.
Figures & tables
Method
Trivia Fantasy
EmbodiedEval
SciKnowEval (L3)
ToolAlpaca
Chem.
Phys.
Biology
Materials
Base model
02.8±1.9
57.3±2.0
31.9±0.3
53.8±0.1
22.1±4.1
68.4±0.8
47.6±1.7
SDPO
61.0±7.1
73.7±1.0
35.9±3.4
54.9±1.9
27.5±0.7
73.2±0.3
59.5±0.5
W2S-OPD
41.7±7.2
60.8±4.2
35.0±1.1
54.8±0.4
27.3±1.9
72.1±0.4
59.7±2.0
Ours
98.3±2.9
82.1±3.0
50.2±2.9
66.9±1.6
28.2±1.8
73.7±1.5
61.9±2.4
Table 1: Our method achieves the highest mean score across all benchmarks and SciKnowEval subjects. Scores (%; higher is better) are means with 95% confidence intervals; the highest mean in each column is bold. Trained-model scores are averaged over three checkpoint evaluations; base-model scores are averaged over independent evaluation passes (Section 5.1 ).
Appendix figures & tables3 assets
Supplementary material from the paper’s appendix.
Appendix
Policy pair
Vocabulary size, horizon
Relative gap
Neural, KL calibration 0.02 nats
32,768,512
2.0%
Neural, KL calibration 0.5 nats
32,768,512
3.3%
Appendix
Table 2: Relative reduction from unskewed Jeffreys for perturbed neural policies. Entries are Monte Carlo estimates using 4,096 rollouts per policy, shared prefixes for both objectives, and exact vocabulary sums.
Trivia Fantasy
EmbodiedEval
SciKnowEval (L3)
ToolAlpaca
Student
Qwen3.5-2B
Qwen3.5-4B
Qwen3.5-4B
Qwen3.5-4B
Training set
20 questions
3,765 scenarios
450–1,890 per subject
4,046 requests
Evaluation set
20 questions
12 suites × 20 tasks
50 / 210 / 80 / 94 (bio. / chem. / phys. / mat.)
68 requests
Learning rate
10−6
10−6
10−7
5×10−7
Warmup steps
5
50
50
50
Sequences per update
512
128
128
128
Appendix
Table 3: Settings shared by all methods on each benchmark. All methods use Adam with weight decay 0.01 , gradient-norm clipping at 1 , a constant learning rate after linear warmup, training sampling at temperature 1 and top- p0.95 , and Gemini 3.7 Flash at temperature 1 as the feedback judge. A rollout becomes stale once the policy that generated it is more than four weight synchronizations old.
Ours
SDPO
W2S-OPD
Teacher
frozen initial student
frozen initial student
frozen initial student
Teacher contexts
with and without feedback
with feedback
with and without feedback
Training rollouts
student and teacher, λ=1/2
student only
student only
Token divergence
α -skew Jeffreys, α=0.01
reverse KL
reverse KL
Target
corrected target equation 6 , β=10
feedback-conditioned teacher
pW2S , contrast strength 0.75 , anchor π0
Importance weights
prefix weights equation 10 , forward weight clipped at cIS=2