Teacher-Student Learning

Momentum

7 papers in the last four weeks, against 2 the four weeks before. 0.1% of all new papers.

Jul 13Week of Sep 28

Latest papers 162

All topics
CardsList
  1. Consistently Informative Soft-Label Temperature for Knowledge Distillation

    May 19, 2026Hoang-Chau Luong, Nghia Van Vo, Kaiqi Zhao +1Teacher-Student LearningKnowledge Distillation

  2. LIFT and PLACE: A Simple, Stable, and Effective Knowledge Distillation Framework for Lightweight Diffusion Models

    May 19, 2026Hyunsoo Han, Sangyeop Yeo, Jaejun YooModel CompressionDiffusion Model Distillation

  3. Decoupling KL and Trajectories: A Unified Perspective for SFT, DAgger, Offline RL, and OPD in LLM Distillation

    May 16, 2026Anhao Zhao, Haoran Xin, Yingqi Fan +3Reasoning DistillationOn-Policy Distillation

  4. PopuLoRA: Co-Evolving LLM Populations for Reasoning Self-Play

    May 16, 2026Roger Creus Castanyer, Geoffrey Bradway, Lorenz Wolf +3Self-Play RLRL for Language Model Reasoning

  5. How to Choose Your Teacher for Fine Grained Image Recognition

    May 15, 2026Oswin Gosal, Edwin Arkel Rios, Augusto Christian Surya +3Fine-Grained Image ClassificationTeacher-Student Learning

  6. DeltaPrompts: Escaping the Zero-Delta Trap in Multimodal Distillation

    May 15, 2026Jaehun Jung, Hyunwoo Kim, Brandon Cui +4VLM DistillationVLM Reasoning

  7. Distribution Corrected Offline Data Distillation for Large Language Models

    May 13, 2026Yumeng Zhang, Zhengbang Yang, Yevin Nikhel Goonatilake +1CoT DistillationEfficient Language Model Reasoning

  8. Prefix Teach, Suffix Fade: Local Teachability Collapse in Strong-to-Weak On-Policy Distillation

    May 13, 2026Kaiyuan Liu, Ziyuan Zhuang, Yang Bai +3Selective Knowledge DistillationLanguage Model Distillation

  9. Teaching and Learning under Deductive Errors

    May 13, 2026Jan Arne Telle, Brigt Håvardstun, Jose Hernandez-OralloParameterized ComplexityPAC Learning

  10. On the Generalization of Knowledge Distillation: An Information-Theoretic View

    May 13, 2026Bingying Li, Haiyun HeInformation-Theoretic Generalization BoundsStatistical Learning Theory

  11. Multi-Rollout On-Policy Distillation via Peer Successes and Failures

    May 12, 2026Weichen Yu, Xiaomin Li, Yizhou Zhao +8RL for Language Model ReasoningLanguage Model Distillation

  12. Anti-Self-Distillation for Reasoning RL via Pointwise Mutual Information

    May 12, 2026Guobin Shen, Xiang Cheng, Chenxiao Zhao +4On-Policy Self-DistillationRL for Language Model Reasoning

  13. Physics-Informed Teacher-Student Ensemble Learning for Traffic State Estimation with a Varying Speed Limit Scenario

    May 11, 2026Archie J. Huang, Dongdong Wang, Shaurya Agarwal +3Ensemble LearningTraffic Flow Estimation

  14. The Many Faces of On-Policy Distillation: Pitfalls, Mechanisms, and Fixes

    May 11, 2026Siqi Zhu, Xuyan Ye, Hongyu Lu +2On-Policy Self-DistillationOn-Policy Distillation

  15. SplitFed-CL: A Split Federated Co-Learning Framework for Medical Image Segmentation with Inaccurate Labels

    May 11, 2026Zahra Hafezi Kafshgari, Hadi Hadizadeh, Parvaneh SaeediSplit Federated LearningNoisy-Label Learning

  16. STEPS: Selective On-Policy Self-Distillation for Reasoning

    May 11, 2026Jiaxuan Wang, Xuan Ouyang, Zhiyu Chen +2On-Policy Self-DistillationRL for Language Model Reasoning

  17. On-Policy Distillation with Best-of-N Teacher Rollout Selection

    May 10, 2026Ke Zhang, Yunjie Tian, Dongdi Zhao +4CoT DistillationBest-of-N Sampling

  18. Cornerstones or Stumbling Blocks? Deciphering the Rock Tokens in On-Policy Distillation

    May 10, 2026Yuxuan Jiang, Runchao Li, Shubhashis Roy Dipta +2Language Model DistillationOn-Policy Distillation

  19. Trajectory as the Teacher: Few-Step Discrete Flow Matching via Energy-Navigated Distillation

    May 8, 2026Amin Karimi Monsefi, Dominic Culver, Nikhil Bhendawade +3Flow MatchingDiffusion Language Model Inference

  20. KL for a KL: On-Policy Distillation with Control Variate Baseline

    May 8, 2026Minjae Oh, Sangjun Song, Gyubin Choi +2Language Model DistillationPolicy Gradient Methods

  21. Prune-OPD: Efficient and Reliable On-Policy Distillation for Long-Horizon Reasoning

    May 8, 2026Zhicheng Yang, Zhijiang Guo, Yifan Song +5RL for Language Model ReasoningLanguage Model Distillation

  22. SimCT: Recovering Lost Supervision for Cross-Tokenizer On-Policy Distillation

    May 8, 2026Jie Sun, Mao Zheng, Mingyang Song +6Cross-Tokenizer Knowledge DistillationLanguage Model Distillation