cs.LGSep 30, 2026

SHARPO: Segment-Level Credit Assignment for Agentic Reinforcement Learning

Authors: Xinchen Du, Zhengze Zhou, Wenhui Zhu, Han Yu, Sen Na, Rohit Jain, Alborz Geramifard

Organizations: LinkedIn Corporation · Georgia Institute of Technology

Abstract

Agentic reinforcement learning (RL) trains a large language model (LLM) to act over long, multi-step interactions. However, a single localized error can cause task failure, while trajectory-level rewards provide limited guidance for assigning credit to individual decisions. To address this limitation, we introduce Segment-level Hindsight Advantage Reweighting for Policy Optimization (SHARPO), a credit-assignment mechanism that refines Group Relative Policy Optimization (GRPO) at the level of environment-facing segments. Inspired by the existing on-policy self-distillation (OPSD) method, SHARPO computes teacher-student log-probability gaps within each segment and uses the resulting signal to compute a bounded multiplier on the GRPO advantage. This multiplier is shared by all tokens within the segment, allowing credit to vary across different segments. With Qwen2.5-7B-Instruct, SHARPO outperforms existing baselines on the ALFWorld and WebShop benchmarks, including GRPO, SDAR, RLSD, and StepOPSD.

Figures & tables

Explore similar work

CardsList
  1. Targeting Pivotal Decisions for Credit Assignment in Agentic Reinforcement Learning

    Sep 28, 2026Dongwon Jung, Hemanth Neelgund Ramesh, Yifan Wang +7Credit AssignmentAgentic Reinforcement Learning

  2. Beyond Trajectory-Level Attribution: Graph-Based Credit Assignment for Agentic Reinforcement Learning

    May 26, 2026Xin Cheng, Shuo He, Lang Feng +4Group-Based Reinforcement LearningAgentic Reinforcement Learning

  3. Semantic Consistency Policy Optimization for Reinforcement Learning of LLM Agents

    Jun 24, 2026Peng Xu, Sijia Chen, Junzhuo Li +1Large Language Model AgentsStep-Level Credit