cs.AISep 28, 2026

UniOPSD: Unifying Outcome and Hindsight Feedback for Agentic Reinforcement Learning

Authors: Zenghuang Fu, Zhaoyang Li, Qiuyuan Ai, Xiaofeng Han, Zelong Zheng, Haoyu Wu, Tianyu Fu, Chenxu Zhao, +3 more

Organizations: University of Chinese Academy of Sciences · Institute of Automation, Chinese Academy of Sciences · Peking University · Mininglamp Technology · Key Laboratory of Computing Power Network and Information Security, Ministry of Education; Shandong Computer Science Center, Qilu University of Technology (Shandong Academy of Sciences) · Key Laboratory of Computing Power Internet and Service Computing, Shandong Fundamental Research Center for Computer Science

Abstract

Reinforcement learning has become an effective approach to training language model agents, but sparse and delayed outcome rewards provide limited guidance for credit assignment across long interaction sequences. Recent work on on-policy self-distillation (OPSD) offers complementary supervision by evaluating a policy's sampled responses under privileged training-time context. However, our diagnostics show that positive average agreement between outcome and hindsight feedback coexists with substantial local disagreement, raising the question of how to allocate influence between them at each decision. We introduce UniOPSD (Unified On-Policy Self-Distillation), which unifies these feedback sources through adaptive local credit arbitration. UniOPSD constructs comparable credit estimates from environmental returns and successful-peer hindsight at shared interaction anchors. Historical agreement determines the global mixing level, while current signal availability and relative precision adjust each source's influence at individual decisions. The episode-level outcome contribution is retained, and bounded token modulation refines the fused step credit for policy optimization. With Qwen2.5-3B-Instruct and Qwen2.5-7B-Instruct, UniOPSD achieves ALFWorld success rates of 82.8%82.8\% and 83.6%83.6\%, WebShop success rates of 75.0%75.0\% and 82.0%82.0\%, and Search-QA aggregate accuracies of 45.3%45.3\% and 49.8%49.8\%, respectively. On 3B WebShop, UniOPSD improves over SDAR by 7.07.0 percentage points. Our code is available at https://github.com/Zenghuang-Fu/Uniopsd

Figures & tables

Appendix figures & tables8 assets

Supplementary material from the paper’s appendix.

Appendix

Explore similar work

CardsList
  1. AgentOPSD: Recursive Self-Distillation for Agentic Reinforcement Learning

    Aug 6, 2026Zi-Han Wang, Zhengxi Lu, Zhiyuan Yao +10Agentic Reinforcement LearningUnsupervised On-Policy Self-Distillation

  2. Agentic Reinforcement Learning with Observation-Calibrated Self-Distillation

    Aug 5, 2026Yi Yang, Cong Qin, Xiaodan Liu +8Unsupervised On-Policy Self-DistillationSelf-Distillation Framework

  3. StepOPSD: Step-Aware Online Preference Distillation for Agent Reinforcement Learning

    May 26, 2026Yanfei Zhang, Xu Lin, Chenglin WuOnline-Policy DistillationRp-Opsd