Know Thyself, Teach Thyself: Internal Information Flow for Selective Self-Distillation
Organizations: Peking University · University of Oxford · The University of Hong Kong
Abstract
Self-distillation turns knowledge distillation into a closed learning loop and offers a path toward recursive self-improvement. Without an external teacher, however, the model must determine both what information can improve its supervision and which induced changes should be learned. Existing methods typically improve teacher-generated data or select training examples in isolation, leaving the information transferred between these stages unmeasured. We introduce InFlow, a retrieval-guided on-policy self-distillation framework that models this process as potential-to-realized information flow. InFlow first retrieves potentially informative sources using certainty-calibrated hidden-state trajectories, then measures their realized effect through the Jensen--Shannon divergence between the teacher's initial and retrieval-conditioned answer beliefs. Examples with larger belief shifts are selected for on-policy distillation. Our analysis formalizes the information optimized by retrieval and selection and relates the answer-level shift to the teacher--student distillation gap. Across four open-weight language models and three knowledge domains, InFlow achieves the strongest cross-model average among the compared selection methods, with ablations supporting both stages of the framework. Our code is available at https://github.com/1240148048/INFLOW.
Figures & tables
| Model | Domain | Original | Full | Redundancy | Consistency | NEURON | Mutual | InFlow |
|---|---|---|---|---|---|---|---|---|
| Qwen3-4B | Math | |||||||
| Natural Sci. | ||||||||
| Hum. & Soc. | ||||||||
| Overall | ||||||||
| Qwen3.5-9B | Math | |||||||
| Natural Sci. |
Appendix figures & tables6 assets
Supplementary material from the paper’s appendix.
Appendix
| Model | Domain | Original | Full | Redundancy | Consistency | NEURON | Mutual | InFlow |
|---|---|---|---|---|---|---|---|---|
| Qwen3-4B | Math | 47.0 | 50.0 | 51.0 | 50.0 | 46.0 | 50.0 | 52.0 |
| Natural Sci. | 78.5 | 79.5 | 78.5 | 78.5 | 80.0 | 79.5 | 79.0 | |
| Hum. & Soc. | 44.5 | 44.0 | 45.0 | 46.0 | 45.0 | 45.5 | 46.0 | |
| Overall | 58.6 | 59.4 | 59.6 | 59.8 | 59.2 | 60.0 | 60.4 | |
| Qwen3.5-9B | Math | 53.0 | 53.0 | 54.0 | 49.0 | 54.0 | 50.0 | 54.0 |
| Natural Sci. | 80.5 | 80.0 | 81.0 | 82.0 | 82.5 | 81.0 | 82.5 |
| JS rank | Fix | Break | W W | R R | Fix Break | Initial Acc. | Teacher Acc. | Mean JS |
|---|---|---|---|---|---|---|---|---|
| 0–10% | 36.0 | 30.0 | 34.0 | 0.0 | 6.0 | 30.0 | 36.0 | 0.6931 |
| 10–20% | 36.0 | 16.0 | 48.0 | 0.0 | 20.0 | 16.0 | 36.0 | 0.5898 |
| 20–30% | 12.0 | 20.0 | 50.0 | 18.0 | 38.0 | 30.0 | 0.3242 | |
| 30–40% | 0.0 | 0.0 | 54.0 | 46.0 | 0.0 | 46.0 | 46.0 | 0.0182 |
| 40–50% | 0.0 | 0.0 | 26.0 | 74.0 | 0.0 | 74.0 | 74.0 | 0.0000 |
| 50–60% | 0.0 | 0.0 | 22.0 | 78.0 | 0.0 | 78.0 | 78.0 | 0.0000 |
| JS rank | Fix | Break | W W | R R | Fix Break | Initial Acc. | Teacher Acc. | Mean JS |
|---|---|---|---|---|---|---|---|---|
| 0–10% | 26.0 | 34.0 | 40.0 | 0.0 | 34.0 | 26.0 | 0.6931 | |
| 10–20% | 26.0 | 20.0 | 54.0 | 0.0 | 6.0 | 20.0 | 26.0 | 0.6414 |
| 20–30% | 26.0 | 16.0 | 56.0 | 2.0 | 10.0 | 18.0 | 28.0 | 0.4223 |
| 30–40% | 4.0 | 4.0 | 68.0 | 24.0 | 0.0 | 28.0 | 28.0 | 0.1815 |
| 40–50% | 0.0 | 0.0 | 62.0 | 38.0 | 0.0 | 38.0 | 38.0 | 0.0461 |
| 50–60% | 0.0 | 0.0 | 38.0 | 62.0 | 0.0 | 62.0 | 62.0 | 0.0000 |
| Method | Retrieval rule | Teacher-data selection rule |
|---|---|---|
| Original | None. | No adaptation. |
| Full | Retrieve the sources with the largest cosine similarity between final-layer answer representations. | Retain all generated teacher data. |
| Redundancy | Use the same final-layer cosine retrieval as Full . | Iteratively remove the example most similar to another retained teacher representation until examples remain. |
| Consistency | No retrieved demonstrations; instead, we obtain more confident training data through multiple sampling, consistency checks, and self-feedback. | Retain the targets with the largest modal-answer agreement among generations. |
| NEURON | Retrieve the sources with the largest Jaccard overlap between sets of positively contributing neurons. | Retain the targets with the smallest mean number of activated contributing neurons across teacher generations. |
| Mutual | Retrieve the sources that provide more conditional mutual information for the target data using a greedy algorithm. | Retain the targets with the largest aggregated retrieval-information score. |