Fisher-Informed Recalibration for Feedback-Based On-Policy Self-Distillation of LLMs
Organizations: Purdue University · Yonsei University · University at Buffalo-SUNY
Abstract
Feedback-based on-policy self-distillation has emerged as a promising approach for enabling foundation models, more specifically Large Language Models (LLMs), to learn from their own outputs under external feedback, with a single model serving as both teacher and student. However, such methods can exhibit unstable optimization, conducive to performance collapse during training. To address this limitation, we propose FIRE (Fisher-Informed REcalibration), a dual-branch framework that recalibrates the supervision applied to correct and incorrect on-policy outputs during fine-tuning. For correct responses, FIRE replaces self-distillation with re-weighted on-policy SFT, while for incorrect ones FIRE identifies feedback components that disproportionately influence the teacher-induced update and recalibrates the feedback-conditioned target accordingly. Both branches are influenced by a token-level radius derived in part from a softmax Fisher trace. FIRE separates which direction feedback should move the model from how far the model should move in that direction, while leaving well-behaved feedback supervision unchanged. Our experiments demonstrate that FIRE provides substantially more stable self-distillation while maintaining strong downstream performance, particularly in settings where standard feedback-conditioned distillation becomes unstable.
Figures & tables
| Model | Algorithm | GSM8K | ASDiv | AQuA-RAT | Avg. |
|---|---|---|---|---|---|
| Qwen2.5-1.5B-Instruct | Full Context | 60.65 | |||
| On-Policy SFT | 59.26 | ||||
| Veto | 58.39 | ||||
| TOP-D | 4.17 | ||||
| TrOPD | 53.76 | ||||
| SRPO | 63.25 |
Appendix figures & tables6 assets
Supplementary material from the paper’s appendix.
Appendix
| Model | Algorithm | Biology | Chemistry | CS | Economics | Engineering | Physics | Overall |
|---|---|---|---|---|---|---|---|---|
| Qwen2.5-14B-Instruct | Full Context | |||||||
| TOP-D | ||||||||
| TrOPD | ||||||||
| SRPO | ||||||||
| DemoPSD | ||||||||
| SCOPE |
| Model | Algorithm | GSM8K | ASDiv | AQuA-RAT | Avg. |
|---|---|---|---|---|---|
| Qwen2.5-1.5B-Instruct | No Attribution | 61.86 | |||
| No Projection | 58.97 | ||||
| Hard Gradient Excision | 59.26 | ||||
| FIRE (Ours) | 63.83 | ||||
| Model | Algorithm | ARC-Challenge | WinoGrande | WiC | Avg. |
| Qwen2.5-3B-Instruct | No Attribution | 72.22 |
| Qwen2.5-1.5B-Instruct | |||
|---|---|---|---|
| LR Multiplier | GSM8K | ASDiv | AQuA-RAT |
| LR ( ): GSM8K/ASDiv ; AQuA-RAT . | |||
| Model | Dataset | Peak LR | Warmup | Steps | Max context | Max new | Train pool | Eval. | ||
| Qwen2.5-1.5B | GSM8K | 0.03 | 25 | 200 | 2560 | 384 | 1024 | 192 | ||
| Qwen2.5-1.5B | ASDiv | 0.03 | 25 | 200 | 2304 | 320 | 1600 | 192 | ||
| Qwen2.5-1.5B | AQuA-RAT | 0.03 | 25 | 250 | 3072 | 512 | 2048 | 192 | ||
| Qwen2.5-3B | ARC-Challenge | 0.01 | 20 | 200 | 2560 | 320 | 2048 | 192 | ||
| Qwen2.5-3B | WinoGrande | 0.01 | 15 | 150 | 2560 | 320 | 2048 | 192 | ||
| Qwen2.5-3B | WiC | 0.01 | 15 | 200 | 2560 | 320 | 2048 | 192 |
| Method | Method-specific settings |
|---|---|
| Full Context | No additional method-specific hyperparameter beyond the shared optimizer, generation, EMA, and LoRA settings. |
| On-Policy SFT | Group size (1.5B) / (3B). |
| Veto | decays linearly from to , forward-KL objective. |
| TOP-D | Group size , proximal coefficient , PPO clipping , one off-policy epoch, one prompt per internal minibatch. |
| TrOPD | Group size , guidance coefficient , maximum guided prefix length . |
| SRPO | Group size for the small-model experiments and for scaling, entropy coefficient , PPO clipping , ratio cap . |
| Method | Qwen 1.5B | Qwen 3B | Qwen 14B |
|---|---|---|---|
| Full Context | |||
| On-Policy SFT | – | ||
| Veto | – | ||
| TOP-D | |||
| TrOPD | |||
| SRPO |