cs.LGSep 28, 2026

GraphHCA: Closed-Form Hindsight Credit Assignment for Long-Horizon LLM Agents

Authors: Haodong Zhu, Yangyang Ren, Changbai Li, Sheng Xu, Linlin Yang, haiguang liu, Baochang Zhang

Organizations: Beihang University · Zhongguancun Academy · Communication University of China · Hangzhou Innovation Institute of Beihang University

Abstract

Group-based reinforcement learning (RL) has advanced large language models (LLMs) and is increasingly extending to agentic tasks, where sparse terminal rewards make step-level credit assignment essential. Existing methods assign credit from what follows an action in sampled rollouts, but do not explicitly capture its retrospective relation to the realized outcome. Hindsight credit assignment (HCA) instead attributes credit through the ratio of hindsight to behavior-policy probabilities, but estimating the hindsight distribution requires an auxiliary model or an extra pass. To address this estimation bottleneck, we propose GraphHCA, a model-free realization of HCA that eliminates explicit hindsight-distribution estimation. For terminal-goal tasks with deterministic transitions, Bayes' rule reduces the hindsight ratio to a ratio of behavior-policy success probabilities at consecutive states. Taking logs yields a state-wise success potential, whose increment across a transition provides step-level credit. GraphHCA estimates this potential from pooled rollouts through a discounted recursion on the induced transition graph, which admits a unique fixed point on any directed graph. The resulting step-level signal is combined with the trajectory-level advantage, requiring neither a learned hindsight model nor an extra forward pass and recovering GRPO when the step-level weight is zero. Among all compared baselines, GraphHCA achieves state-of-the-art results on ALFWorld and WebShop at both LLM scales, and on Sokoban with a vision-language agent. For example, on ALFWorld it improves overall success rate by up to 24.6 points over GRPO and by up to 4.7 points over the strongest step-level baseline.

Figures & tables

Appendix figures & tables6 assets

Supplementary material from the paper’s appendix.

Appendix

Explore similar work

CardsList
  1. Beyond Trajectory-Level Attribution: Graph-Based Credit Assignment for Agentic Reinforcement Learning

    May 26, 2026Xin Cheng, Shuo He, Lang Feng +4Group-Based Reinforcement LearningAgentic Reinforcement Learning

  2. Granularity-Adaptive Credit Assignment for Long-Horizon LLM Agent Reinforcement Learning

    Sep 14, 2026Taoran Liang, Yang Liu, Shang Luo +9Credit AssignmentLarge Language Model Agents

  3. Back to the Definition: Estimating Step-Level Advantages via Trajectory Graphs for Agentic Reinforcement Learning

    Sep 24, 2026Xincheng Yao, Haobo Fu, Weiming Liu +1Group-Based Reinforcement LearningAgentic Reinforcement Learning