cs.AISep 27, 2026

SeOPD: Self-Evolving LLMs via Online Policy Distillation from Self-Generated Chain-of-Thought

Authors: Xiaoshu Chen, Sihang Zhou, Ke Liang, Xinwang Liu

Organizations: National University of Defense Technology

Abstract

Recent advances in online policy self-distillation (OPSD) have demonstrated that large language models (LLMs) can improve their capabilities by leveraging external privileged information (PI), such as manual annotations or feedback from external environments. However, obtaining accurate annotations and constructing sophisticated environments often require substantial human effort and computation, limiting the scalability of OPSD. While a few recent studies have explored self-improvement without external PI, the resulting gains remain limited. In this work, we explore whether LLMs can achieve comparable self-improvement without external PI. Our key observation is that a single LLM can support multiple reasoning modes, such as deep-thinking and non-thinking modes, with deep thinking generating additional information during reasoning. Based on this observation, we propose Self-Evolving Online Policy Distillation (SeOPD), which enables LLMs to distill and internalize information generated by their own chain of thought (CoT). Specifically, it (1) generates CoT with the deep-thinking mode, (2) produces responses with the non-thinking mode, and (3) uses the generated CoT as PI to provide token-level supervision for the non-thinking response, allowing new information inferred during reasoning to guide the non-thinking mode and be internalized into the shared model parameters, thereby improving both non-thinking and deep-thinking capabilities. Extensive experiments across LLMs and tasks demonstrate the effectiveness of SeOPD.

Figures & tables

Appendix figures & tables5 assets

Supplementary material from the paper’s appendix.

Appendix

Explore similar work

CardsList
  1. On-Policy Self-Distillation without Any Supervision

    Aug 6, 2026Yijiang Li, Bingyang Wang, Yijun Liang +3Unsupervised On-Policy Self-DistillationSelf-Distillation Framework

  2. Purified OPSD: On-Policy Self-Distillation Without Losing How to Think

    Jul 2, 2026Zhanming Shen, Jintao Tong, Shaotian Yan +9Unsupervised On-Policy Self-DistillationTeacher

  3. DemoPSD: Disagreement-Modulated Policy Self-Distillation

    Jul 2, 2026Yunhe Li, Hao Shi, Wenhao Liu +5Unsupervised On-Policy Self-DistillationTeacher-Student Distillation