We study training LLM judges from natural language feedback, especially for subjective tasks where the verdict depends strongly on which evaluation criteria the judge invokes and how it weighs them. The dominant approach, outcome-supervised RL (e.g., GRPO), credits every token in the rollout with a single scalar determined only by the accuracy of the final verdict, providing no separate credit at the criterion-choice tokens and ignoring the rich language feedback (e.g., preference rationales) that naturally accompanies preference labels. Self-Distillation (SD) is one natural way to use this language feedback: the same model, conditioned on this feedback, acts as a teacher providing dense, position-level supervision. However, not all positions carry equally useful signal. Using the per-position entropy shift between teacher and student, we identify two regimes: context sharpening, where the teacher concentrates probability on a particular feedback-aligned criterion expression, and context spreading, where the teacher distributes probability across multiple feedback-aligned alternatives. We interpret these patterns as follows: sharpening encourages memorization of a particular criterion expression, whereas spreading promotes semantic understanding by preserving these alternatives. Motivated by this asymmetry, we introduce position masking based on the entropy shift that retains the lower tail of the entropy-shift distribution. Experiments show that masking higher-entropy-shift positions improves out-of-distribution generalization over naive SD. The resulting self-distilled judges outperform judges trained with outcome-supervised RL by 2-9 percentage points on the evaluated subjective subcategories, while remaining competitive on objective ones.
Figures & tables
Figure 2 : Both ΔHt tails account for most of the distillation loss. (a) Three-way per-sequence percentile partition (top-30% ΔHt / middle 40% / bottom-30%): both tails carry ∼ 20–28 × the per-position KL of the neutral middle. (b) Binned mean reverse KL along ΔHt using 20 equal-count bins (95% x-range shown).
Figure 3 : Context sharpening on two HelpSteer3-Preference rollouts (Qwen3-30B-A3B-Instruct-2507). Each panel shows the user task, the rationale (anchor word red ), the student rollout excerpt with the annotated position ⟨ Pos N⟩ marked in blue , and student vs. teacher top-5 next-token distributions at that position. All positions sit at criterion-naming slots in the rollout’s opening evaluation-criteria list. The teacher concentrates probability on a token semantically equivalent to the rationale’s anchor while the student is uncertain across multiple plausible criteria. Two further rollouts are shown in Figure 6 .
Figure 4 : Context spreading on two HelpSteer3-Preference rollouts (Qwen3-30B-A3B-Instruct-2507). Same layout as Figure 3 . Each panel features a position where the student commits to a criterion that is non-decisive in the rationale (top-1 probability ≥0.79 ), while the teacher reopens with a spread of alternatives that include semantic equivalents of the rationale’s decisive criterion. All positions sit at criterion-naming slots in the rollout’s opening evaluation-criteria list. Two further rollouts are shown in Figure 7 .
Figure 5 : Distinct semantic clusters of evaluation criteria invoked by trained 30B judges across a sweep of complete-link cosine clustering thresholds. Naive SD invokes a markedly smaller pool than SD+mask at every threshold; the gap widens with stricter clustering.
Models
RM-Bench
RewardBench v2
Total Avg.
Chat
Code
Math
Safety
Overall
Factuality
Focus
Math
Precise IF
Safety
Overall
Prompting
Claude-Sonnet-4-20250514
79.93
81.77
93.05
95.34
87.52
84.00
87.27
85.25
56.88
94.00
81.48
84.50
DeepSeek-R1
78.38
79.29
94.01
92.27
85.99
74.53
80.40
87.43
43.12
86.00
74.30
80.15
Training (Small Model < 10B)
Think-RM-8B
65.20
54.63
71.92
91.71
70.87
49.26
62.42
58.47
26.88
75.56
54.52
62.70
Table 1 : Judge accuracy (%) on RM-Bench and RewardBench v2 . Dark gray (in bold) and light gray highlight the best and second-best performance per column, respectively. RationaleRM-30B numbers are taken from Wang et al. [2026a] as the model has not been released; RewardBench v2 cells are left empty.
Selector
Creative-writing win rate
Dr. GRPO
0.556
Naive SD
0.612
SD+mask ( ρ=0.7 )
0.632
Table 2 : Creative-writing win rate of responses selected through multi-round best-of-eight tournaments. Each tournament uses a different Qwen3-30B-A3B-Instruct judge as its pairwise selector; selected responses are evaluated against the Arena-Hard-v2 reference responses by Claude Sonnet 5.
ρ (fraction masked)
RM-Bench
RewardBench v2
Total Avg.
Chat
Code
Math
Safety
Overall
Factuality
Focus
Math
Precise IF
Safety
Overall
0.00 (naive SD)
76.74
76.56
94.45
94.26
85.50
78.53
86.26
81.42
40.62
87.33
74.83
80.17
0.30
77.52
75.78
95.02
94.21
85.63
80.00
83.03
85.79
43.75
85.56
75.63
80.63
0.50
77.17
77.49
94.85
94.15
85.92
78.95
83.43
84.15
46.88
83.78
75.44
80.68
0.70 (canonical)
80.53
80.07
95.63
94.81
87.76
79.16
85.66
84.15
48.75
88.44
77.23
82.50
Table 3 : ρ -sweep on Qwen3-30B-A3B-Instruct, all masking the top- ρ fraction by ΔHt and training on the remaining bottom (1−ρ) .
Selector
RM-Bench
RewardBench v2
Total Avg.
Chat
Code
Math
Safety
Overall
Factuality
Focus
Math
Precise IF
Safety
Overall
Mask top ρ fraction by ΔHt (ours; drops highest-shift positions)
Table 4 : Selector ablation on Qwen3-30B-A3B-Instruct, all at matched mask fraction ρ=0.7 .
Appendix figures & tables5 assets
Supplementary material from the paper’s appendix.
Appendix
Figure 6 : Additional context-sharpening examples on two HelpSteer3-Preference rollouts (Qwen3-30B-A3B-Instruct-2507).
Figure 7 : Additional context-spreading examples on two HelpSteer3-Preference rollouts (Qwen3-30B-A3B-Instruct-2507).
Models
RM-Bench
RewardBench v2
Total Avg.
Chat
Code
Math
Safety
Overall
Factuality
Focus
Math
Precise IF
Safety
Overall
Qwen3-4B-Instruct + Dr. GRPO
68.91
73.83
94.01
90.93
81.92
62.74
75.76
86.34
43.12
83.78
70.35
76.14
Qwen3-4B-Instruct + Naive SD (rubric)
75.45
75.19
93.57
92.34
84.14
61.68
74.23
84.39
36.00
79.09
67.08
75.61
Qwen3-4B-Instruct + SD+mask (rubric)
77.26
75.97
93.28
91.91
84.61
65.89
75.05
84.97
48.00
79.55
70.69
77.65
Qwen3-4B-Instruct + Naive SD (rationale)
75.71
73.34
93.51
92.06
83.66
68.42
80.20
85.79
38.75
85.56
71.74
77.70
Qwen3-4B-Instruct + SD+mask (rationale)
77.95
75.88
94.10
90.95
84.72
70.74
77.78
84.15
45.00
89.11
73.36
79.04
Appendix
Table 5 : Judge accuracy (%) for Qwen3-4B-Instruct using preference rationales or generated per-example rubrics as language feedback. Dark gray (in bold) and light gray highlight the best and second-best performance per column, respectively.
ρ (fraction masked)
RM-Bench
RewardBench v2
Total Avg.
Chat
Code
Math
Safety
Overall
Factuality
Focus
Math
Precise IF
Safety
Overall
0.00 (naive SD)
75.71
73.34
93.51
92.06
83.66
68.42
80.20
85.79
38.75
85.56
71.74
77.70
0.30
75.54
74.51
93.97
91.38
83.85
67.79
75.96
83.06
41.88
87.33
71.20
77.53
0.50
76.92
75.54
93.82
91.74
84.50
67.58
76.36
83.61
46.88
88.22
72.53
78.52
0.70 (canonical)
77.95
75.88
94.10
90.95
84.72
70.74
77.78
84.15
45.00
89.11
73.36
79.04
Appendix
Table 6 : ρ -sweep on Qwen3-4B-Instruct, all masking the top- ρ fraction by ΔHt and training on the remaining bottom (1−ρ) .
Selector
RM-Bench
RewardBench v2
Total Avg.
Chat
Code
Math
Safety
Overall
Factuality
Focus
Math
Precise IF
Safety
Overall
Mask top ρ fraction by ΔHt (ours; drops highest-shift positions)
Conditioning a language model on additional context, such as feedback on a previous attempt, typically improves its response. Self-distillation trains the model to retain this improvement when the context is not present. The method works by matching the model's output distribution under two settings: a student that sees only the question, and a self-teacher that also sees the context. What the model learns therefore depends on what context the self-teacher receives, yet the design of this context remains largely unexplored. We study context design for self-distillation by training a solver on feedback from a frozen critic. We compare three conditions: (i) a binary reward (GRPO), (ii) the reference solution, and (iii) a step-by-step critique aligned to the solver's reasoning trace. Step-aligned critique yields the largest gains, outperforming GRPO by 16.11 points and reference-solution-conditioned self-distillation by 5.27 points (Avg@12). Per-token advantage analysis reveals why: step-aligned feedback targets only the tokens where reasoning fails, leaving correct behavior intact. Conditioning on the reference solution, by contrast, pressures the model to change its behavior at every token (even correct steps) because an alternative derivation inevitably differs in phrasing and approach. This suggests that structural alignment between feedback and the solver's reasoning is a key driver of self-distillation effectiveness.
Reinforcement learning with verifiable rewards (RLVR) provides reliable outcome supervision for language model reasoning, but a scalar trajectory reward offers limited token-level guidance. Existing self-distillation methods add a privileged teacher but typically assign it a fixed role: direct distribution matching may destabilize successful behavior, while magnitude-only modulation offers little corrective guidance after failure. We observe that successful and failed trajectories require different forms of hindsight supervision. A successful response already contains a valid student-generated reasoning path and can therefore serve as privileged context rather than being replaced by an external rationale. A failed response, however, requires corrective reference information. We introduce Hybrid Hindsight Self-Distillation (H2SD), which jointly adapts teacher context and update strategy to trajectory correctness. For successful trajectories, we construct the teacher context from the verified response and a rephrasing instruction, and use the teacher only to re-evaluate the original response tokens. The resulting probabilities refine token credit assignment without changing the direction determined by the reward. For failed trajectories, a verified reference hint provides corrective guidance through reverse-KL distillation. Experiments on challenging reasoning benchmarks show that H2SD achieves the strongest overall performance among representative RLVR and self-distillation baselines, with stable optimization and a favorable accuracy-efficiency trade-off.
Qiye Cai, Yichuan Ma, Peiji Li +6
1Shanghai Artificial Intelligence Laboratory · 2Harbin Institute of Technology · 3Fudan University +1
Reinforcement learning from verifiable rewards assigns a single scalar to each rollout, leaving token-level credit assignment underspecified in long reasoning traces. On-policy self-distillation addresses this by letting the same model act as a teacher conditioned on privileged information, producing a dense per-token signal. But the common choice of a ground-truth answer is only an endpoint cue: on terse-answer tasks, the teacher falls silent at the intermediate positions where path-level guidance matters most. We propose Hindsight Self-Distillation (HSD), which conditions the teacher on a successful peer rollout drawn from the current training group. Such a peer is an exact sample from the success-conditioned policy, requiring no additional sampled rollouts. By providing a full successful continuation rather than only the final answer, the resulting credit signal concentrates at the divergence position between a failed rollout and a successful peer. Across Qwen3-8B and Qwen3-32B on math and code benchmarks, HSD obtains the best result against GRPO variants and on-policy distillation baselines, with the largest gains on terse-answer tasks such as AIME.
Yu Li, Shu Hong, Tian Lan
Department of Electrical and Computer Engineering, George Washington University