cs.LGSep 27, 2026

Beyond Timestamps: Decision-Aligned On-Policy Distillation for Long-Horizon Agents

Authors: Mingju Chen, Can Lv, Jinrong Liu, Huan Zhang, Heng Chang, Shiji Zhou

Organizations: Beijing Advanced Innovation Center for Future Blockchain and Privacy Computing, Beihang University · School of Artificial Intelligence, Beihang University · Tsinghua University

Abstract

Reinforcement learning with verifiable rewards (RLVR) often relies on sparse outcome rewards, providing coarse supervision for long-horizon agents. On-policy self-distillation (OPSD) complements this signal with dense privileged feedback. However, we identify \emph{Decision--Timestamp Mismatch}: privileged guidance may be misaligned with the student's functional decision because the corresponding decision can occur at a different timestep, while the student's decision itself may span multiple timesteps rather than being tied to a single timestamp. Thus, timestamp-local supervision can misalign both the context and the temporal scope of credit. To address this mismatch, we introduce \textsc{AlignOPSD}, following the principle of aligning supervision before assigning credit. Decision-Aligned Supervision Rectification re-scores the same student-sampled response in functionally matched contexts across sibling rollouts to calibrate local teacher evidence. Semi-Markov Hierarchical Credit Assignment then derives variable-duration decision spans from correspondence changes and uses rectified evidence to allocate outcome-grounded credit across spans and their constituent turns. We evaluate \textsc{AlignOPSD} with Qwen2.5-3B and Qwen2.5-7B on ALFWorld, WebShop, and Search-QA against representative baselines. \textsc{AlignOPSD} outperforms both GRPO and StepOPSD across all eight backbone--aggregate-metric comparisons, improving on GRPO by 5.5--8.7 % and ranking first in six. Additional analyzes examine the two alignment stages and hyperparameter sensitivity between tasks. Our code is avaliable at https://github.com/mingju-c/Align-OPSD

Figures & tables

Appendix figures & tables12 assets

Supplementary material from the paper’s appendix.

Appendix

Explore similar work

CardsList
  1. Dense Is Not Enough: Hierarchical Supervision Allocation for Long-Horizon On-Policy Distillation

    Sep 27, 2026Yuhao Sun, Binrui Wu, Zhuoer Xu +5Efficient On-Policy DistillationTeacher

  2. AgentOPSD: Recursive Self-Distillation for Agentic Reinforcement Learning

    Aug 6, 2026Zi-Han Wang, Zhengxi Lu, Zhiyuan Yao +10Agentic Reinforcement LearningUnsupervised On-Policy Self-Distillation

  3. StepOPSD: Step-Aware Online Preference Distillation for Agent Reinforcement Learning

    May 26, 2026Yanfei Zhang, Xu Lin, Chenglin WuOnline-Policy DistillationRp-Opsd