cs.CLSep 30, 2026

Training LLM Judges from Language Feedback via Position-Selective Self-Distillation

Authors: Ilgee Hong, Changlong Yu, Zhenghao Xu, Xin Liu, Yuwei Zhang, Qin Lu, Bing Yin, Tuo Zhao

Organizations: Georgia Institute of Technology · Amazon · UC San Diego

Abstract

We study training LLM judges from natural language feedback, especially for subjective tasks where the verdict depends strongly on which evaluation criteria the judge invokes and how it weighs them. The dominant approach, outcome-supervised RL (e.g., GRPO), credits every token in the rollout with a single scalar determined only by the accuracy of the final verdict, providing no separate credit at the criterion-choice tokens and ignoring the rich language feedback (e.g., preference rationales) that naturally accompanies preference labels. Self-Distillation (SD) is one natural way to use this language feedback: the same model, conditioned on this feedback, acts as a teacher providing dense, position-level supervision. However, not all positions carry equally useful signal. Using the per-position entropy shift between teacher and student, we identify two regimes: context sharpening, where the teacher concentrates probability on a particular feedback-aligned criterion expression, and context spreading, where the teacher distributes probability across multiple feedback-aligned alternatives. We interpret these patterns as follows: sharpening encourages memorization of a particular criterion expression, whereas spreading promotes semantic understanding by preserving these alternatives. Motivated by this asymmetry, we introduce position masking based on the entropy shift that retains the lower tail of the entropy-shift distribution. Experiments show that masking higher-entropy-shift positions improves out-of-distribution generalization over naive SD. The resulting self-distilled judges outperform judges trained with outcome-supervised RL by 2-9 percentage points on the evaluated subjective subcategories, while remaining competitive on objective ones.

Figures & tables

Appendix figures & tables5 assets

Supplementary material from the paper’s appendix.

Appendix

Explore similar work

CardsList
  1. H2^2SD: Hybrid Hindsight Self-Distillation

    Jul 21, 2026Qiye Cai, Yichuan Ma, Peiji Li +6Reinforcement Learning With Verifiable RewardHindsight

  2. Localizing Credit at the Divergence: Path-Conditioned Self-Distillation for LLM Reasoning

    Jun 14, 2026Yu Li, Shu Hong, Tian LanUnsupervised On-Policy Self-DistillationOnline-Policy Distillation