cs.AISep 24, 2026

Back to the Definition: Estimating Step-Level Advantages via Trajectory Graphs for Agentic Reinforcement Learning

Authors: Xincheng Yao, Haobo Fu, Weiming Liu, Chongyang Zhang

Organizations: School of Information Science and Electronic Engineering, Shanghai Jiao Tong University. · Tencent AI Platform Department. · MoE Key Lab of Artificial Intelligence, AI Institute, Shanghai Jiao Tong University.

Abstract

Group-based reinforcement learning (RL) methods, such as GRPO and its variants, have become a leading paradigm for training reasoning and agentic large language models (LLMs). While their group-normalized advantage estimation is reliable at the response level, it becomes systematically biased at the step level, since coarse-grained trajectory-level advantages are hard to accurately reflect the contribution of individual steps (i.e, failed trajectories may contain valuable steps). Revisiting the foundational RL definition, we notice that GRPO's success on single-turn tasks stems from its advantage estimation strategy, which adheres to the basic definition: the mean reward of multiple actions sampled from the same state constitutes a credible state-value estimate. Extending the faithful estimation to step-level would in principle demand sampling multiple actions from each intermediate state, which is too costly on a per-state basis. To mitigate this issue, we propose a Graph-based Faithful sTep-level credit-assignment framework (GRAFT) that grafts all rollout trajectories into a trajectory graph, recovering node state-values via Bellman iteration on the graph, and assigning credit to each edge by the node value difference. Theoretically, the estimated step-level advantage faithfully adheres to the basic advantage definition in RL. To further ensure the reliability of step-level advantage estimation, we further propose Graph GAE, which extends GAE to the trajectory graph for reducing the impact of state-value estimation bias. Experiments across a range of multi-turn agentic benchmarks show consistent gains over GRPO and superior performance compared to recent agentic RL algorithms. Code will be available at https://github.com/xcyao00/GRAFT.

Figures & tables

Appendix figures & tables9 assets

Supplementary material from the paper’s appendix.

Appendix

Explore similar work

CardsList
  1. Beyond Trajectory-Level Attribution: Graph-Based Credit Assignment for Agentic Reinforcement Learning

    May 26, 2026Xin Cheng, Shuo He, Lang Feng +4Group-Based Reinforcement LearningAgentic Reinforcement Learning

  2. GraphHCA: Closed-Form Hindsight Credit Assignment for Long-Horizon LLM Agents

    Sep 28, 2026Haodong Zhu, Yangyang Ren, Changbai Li +4Group-Based Reinforcement LearningCredit Assignment

  3. Group-Graph Policy Optimization for Long-Horizon Agentic Reinforcement Learning

    Jun 22, 2026Yunan Wang, Minghui Song, Zihan Zhang +6Group-Based Reinforcement Learning