We study training LLM judges from natural language feedback, especially for subjective tasks where the verdict depends strongly on which evaluation criteria the judge invokes and how it weighs them. The dominant approach, outcome-supervised RL (e.g., GRPO), credits every token in the rollout with a single scalar determined only by the accuracy of the final verdict, providing no separate credit at the criterion-choice tokens and ignoring the rich language feedback (e.g., preference rationales) that naturally accompanies preference labels. Self-Distillation (SD) is one natural way to use this language feedback: the same model, conditioned on this feedback, acts as a teacher providing dense, position-level supervision. However, not all positions carry equally useful signal. Using the per-position entropy shift between teacher and student, we identify two regimes: context sharpening, where the teacher concentrates probability on a particular feedback-aligned criterion expression, and context spreading, where the teacher distributes probability across multiple feedback-aligned alternatives. We interpret these patterns as follows: sharpening encourages memorization of a particular criterion expression, whereas spreading promotes semantic understanding by preserving these alternatives. Motivated by this asymmetry, we introduce position masking based on the entropy shift that retains the lower tail of the entropy-shift distribution. Experiments show that masking higher-entropy-shift positions improves out-of-distribution generalization over naive SD. The resulting self-distilled judges outperform judges trained with outcome-supervised RL by 2-9 percentage points on the evaluated subjective subcategories, while remaining competitive on objective ones.
Figures & tables
Figure 2 : Both ΔHt tails account for most of the distillation loss. (a) Three-way per-sequence percentile partition (top-30% ΔHt / middle 40% / bottom-30%): both tails carry ∼ 20–28 × the per-position KL of the neutral middle. (b) Binned mean reverse KL along ΔHt using 20 equal-count bins (95% x-range shown).
Figure 3 : Context sharpening on two HelpSteer3-Preference rollouts (Qwen3-30B-A3B-Instruct-2507). Each panel shows the user task, the rationale (anchor word red ), the student rollout excerpt with the annotated position ⟨ Pos N⟩ marked in blue , and student vs. teacher top-5 next-token distributions at that position. All positions sit at criterion-naming slots in the rollout’s opening evaluation-criteria list. The teacher concentrates probability on a token semantically equivalent to the rationale’s anchor while the student is uncertain across multiple plausible criteria. Two further rollouts are shown in Figure 6 .
Figure 4 : Context spreading on two HelpSteer3-Preference rollouts (Qwen3-30B-A3B-Instruct-2507). Same layout as Figure 3 . Each panel features a position where the student commits to a criterion that is non-decisive in the rationale (top-1 probability ≥0.79 ), while the teacher reopens with a spread of alternatives that include semantic equivalents of the rationale’s decisive criterion. All positions sit at criterion-naming slots in the rollout’s opening evaluation-criteria list. Two further rollouts are shown in Figure 7 .
Figure 5 : Distinct semantic clusters of evaluation criteria invoked by trained 30B judges across a sweep of complete-link cosine clustering thresholds. Naive SD invokes a markedly smaller pool than SD+mask at every threshold; the gap widens with stricter clustering.
Models
RM-Bench
RewardBench v2
Total Avg.
Chat
Code
Math
Safety
Overall
Factuality
Focus
Math
Precise IF
Safety
Overall
Prompting
Claude-Sonnet-4-20250514
79.93
81.77
93.05
95.34
87.52
84.00
87.27
85.25
56.88
94.00
81.48
84.50
DeepSeek-R1
78.38
79.29
94.01
92.27
85.99
74.53
80.40
87.43
43.12
86.00
74.30
80.15
Training (Small Model < 10B)
Think-RM-8B
65.20
54.63
71.92
91.71
70.87
49.26
62.42
58.47
26.88
75.56
54.52
62.70
Table 1 : Judge accuracy (%) on RM-Bench and RewardBench v2 . Dark gray (in bold) and light gray highlight the best and second-best performance per column, respectively. RationaleRM-30B numbers are taken from Wang et al. [2026a] as the model has not been released; RewardBench v2 cells are left empty.
Selector
Creative-writing win rate
Dr. GRPO
0.556
Naive SD
0.612
SD+mask ( ρ=0.7 )
0.632
Table 2 : Creative-writing win rate of responses selected through multi-round best-of-eight tournaments. Each tournament uses a different Qwen3-30B-A3B-Instruct judge as its pairwise selector; selected responses are evaluated against the Arena-Hard-v2 reference responses by Claude Sonnet 5.
ρ (fraction masked)
RM-Bench
RewardBench v2
Total Avg.
Chat
Code
Math
Safety
Overall
Factuality
Focus
Math
Precise IF
Safety
Overall
0.00 (naive SD)
76.74
76.56
94.45
94.26
85.50
78.53
86.26
81.42
40.62
87.33
74.83
80.17
0.30
77.52
75.78
95.02
94.21
85.63
80.00
83.03
85.79
43.75
85.56
75.63
80.63
0.50
77.17
77.49
94.85
94.15
85.92
78.95
83.43
84.15
46.88
83.78
75.44
80.68
0.70 (canonical)
80.53
80.07
95.63
94.81
87.76
79.16
85.66
84.15
48.75
88.44
77.23
82.50
Table 3 : ρ -sweep on Qwen3-30B-A3B-Instruct, all masking the top- ρ fraction by ΔHt and training on the remaining bottom (1−ρ) .
Selector
RM-Bench
RewardBench v2
Total Avg.
Chat
Code
Math
Safety
Overall
Factuality
Focus
Math
Precise IF
Safety
Overall
Mask top ρ fraction by ΔHt (ours; drops highest-shift positions)
Table 4 : Selector ablation on Qwen3-30B-A3B-Instruct, all at matched mask fraction ρ=0.7 .
Appendix figures & tables5 assets
Supplementary material from the paper’s appendix.
Appendix
Figure 6 : Additional context-sharpening examples on two HelpSteer3-Preference rollouts (Qwen3-30B-A3B-Instruct-2507).
Figure 7 : Additional context-spreading examples on two HelpSteer3-Preference rollouts (Qwen3-30B-A3B-Instruct-2507).
Models
RM-Bench
RewardBench v2
Total Avg.
Chat
Code
Math
Safety
Overall
Factuality
Focus
Math
Precise IF
Safety
Overall
Qwen3-4B-Instruct + Dr. GRPO
68.91
73.83
94.01
90.93
81.92
62.74
75.76
86.34
43.12
83.78
70.35
76.14
Qwen3-4B-Instruct + Naive SD (rubric)
75.45
75.19
93.57
92.34
84.14
61.68
74.23
84.39
36.00
79.09
67.08
75.61
Qwen3-4B-Instruct + SD+mask (rubric)
77.26
75.97
93.28
91.91
84.61
65.89
75.05
84.97
48.00
79.55
70.69
77.65
Qwen3-4B-Instruct + Naive SD (rationale)
75.71
73.34
93.51
92.06
83.66
68.42
80.20
85.79
38.75
85.56
71.74
77.70
Qwen3-4B-Instruct + SD+mask (rationale)
77.95
75.88
94.10
90.95
84.72
70.74
77.78
84.15
45.00
89.11
73.36
79.04
Appendix
Table 5 : Judge accuracy (%) for Qwen3-4B-Instruct using preference rationales or generated per-example rubrics as language feedback. Dark gray (in bold) and light gray highlight the best and second-best performance per column, respectively.
ρ (fraction masked)
RM-Bench
RewardBench v2
Total Avg.
Chat
Code
Math
Safety
Overall
Factuality
Focus
Math
Precise IF
Safety
Overall
0.00 (naive SD)
75.71
73.34
93.51
92.06
83.66
68.42
80.20
85.79
38.75
85.56
71.74
77.70
0.30
75.54
74.51
93.97
91.38
83.85
67.79
75.96
83.06
41.88
87.33
71.20
77.53
0.50
76.92
75.54
93.82
91.74
84.50
67.58
76.36
83.61
46.88
88.22
72.53
78.52
0.70 (canonical)
77.95
75.88
94.10
90.95
84.72
70.74
77.78
84.15
45.00
89.11
73.36
79.04
Appendix
Table 6 : ρ -sweep on Qwen3-4B-Instruct, all masking the top- ρ fraction by ΔHt and training on the remaining bottom (1−ρ) .
Selector
RM-Bench
RewardBench v2
Total Avg.
Chat
Code
Math
Safety
Overall
Factuality
Focus
Math
Precise IF
Safety
Overall
Mask top ρ fraction by ΔHt (ours; drops highest-shift positions)