Group Relative Policy Optimization (GRPO) avoids a separate critic by estimating advantages from rollout groups. For multi-turn agents, however, trajectory-level supervision provides coarse, noisy credit: terminal rewards do not locate errors and can penalize useful actions alongside mistakes. Group-in-Group Policy Optimization (GiGPO) and subsequent methods refine supervision through state-conditioned comparisons, but their credit estimates remain sensitive to downstream decisions and outcomes. We introduce RELACE, Retrospective Likelihood-based Action, a critic-free framework that integrates retrospective action assessment with state-conditioned advantage estimation. RELACE evaluates executed actions through teacher-forced likelihood scoring under both their original contexts and outcome-augmented contexts. Comparing these likelihoods yields a trajectory-normalized retrospective factor that captures outcome-dependent changes in action plausibility, rather than hindsight plausibility alone. We use this factor to reweight discounted task returns and construct local advantages by comparing weighted returns among actions from equivalent states within a task. This couples retrospective relevance with observed reward, producing fine-grained credit that complements trajectory-level GRPO supervision. Temporal smoothing and success-protecting masking further stabilize the local signal. RELACE requires neither auxiliary value nor reward models nor additional autoregressive rollouts for credit estimation. Experiments on ALFWorld and WebShop with Qwen2.5-1.5B-Instruct and Qwen2.5-7B-Instruct demonstrate substantial improvements over GRPO, GiGPO, and HCAPO. With the 1.5B model, RELACE achieves 96.35% success on ALFWorld and 79.43% on WebShop, surpassing GiGPO by 5.47 and 5.60 percentage points, respectively.
Figures & tables
Figure 1: Overview of RELACE . Multiple trajectories are sampled for the same task, with colors indicating matched or similar environment states across rollouts. After the terminal outcome is observed, the frozen behavior policy scores the same executed actions under original and retrospective contexts using teacher forcing. The resulting hindsight reweighted returns are compared across matched states to construct the step-level credit signal. Arrow widths illustrate relative retrospective weights.
Figure 2: Validation success over steps 0 – 200 for RELACE (solid) and GiGPO (dashed), using Qwen2.5-1.5B/7B-Instruct. Lines show seed means; shading shows one sample standard deviation, omitted for single-seed baselines. Seed coverage and the † run-assignment note appear in Appendix C .
Variant
Focal target
Baseline input
Reported SR (%)
GiGPO
Q
Q
73.83
Hindsight-score weighting
pHQ
pHQ
75.78
Focal-only ratio weighting
ρQ
Q
60.94∗
RELACE
ρQ
ρQ
79.43
Table 2: Component comparison on WebShop. Columns show the focal target and group-baseline input before smoothing.
Figure 3: RELACE scoring dynamics on WebShop (Qwen2.5-1.5B-Instruct, steps 1 – 200 ): (a) hindsight weight, with a dashed reference at one; (b) hindsight action score; and (c) action score.
Appendix figures & tables1 asset
Supplementary material from the paper’s appendix.
Appendix
Figure 4: Mean training latency on WebShop over logged steps 1 – 200 of one timing run. Labels show seconds per iteration and the share of full logged iteration time. Hindsight scoring (orange) averages 6.85 s, or 7.97% of the full iteration. “Other logged time” includes periodic evaluation, checkpointing, and residual work.
Long-horizon language-model agents trained with reinforcement learning oftenreceive sparse outcome rewards that do not reveal which decisions along a tra-jectory deserve credit. Episode-level advantages provide coarse trajectory-widecredit, while step-level comparisons offer finer resolution with context-dependentestimation noise. We propose Granularity-Adaptive Credit Assignment (GACA),a critic-free method that adaptively mixes episode- and step-level credit for eachdecision during policy optimization. GACA normalizes the sampled response'smean per-token negative log-likelihood (NLL) within each trajectory and uses theresulting criticality score to determine the step-specific mixture. The computationreuses rollout log-probabilities without additional training rollouts or model eval-uations. Our analysis characterizes optimal score-dependent mixing and derivesconditions linking expected NLL to a lower bound on the preferred step-levelweight. Across ALFWorld and WebShop with 1.5B and 7B backbones, GACAachieves the highest reported mean success rates among the compared methods,while introducing negligible additional computation.
Taoran Liang, Yang Liu, Shang Luo +10
Nankai University · Peking University · Supply Chain Tech Team Y, JD.com +2
Group-based reinforcement learning (RL) has advanced large language models (LLMs) and is increasingly extending to agentic tasks, where sparse terminal rewards make step-level credit assignment essential. Existing methods assign credit from what follows an action in sampled rollouts, but do not explicitly capture its retrospective relation to the realized outcome. Hindsight credit assignment (HCA) instead attributes credit through the ratio of hindsight to behavior-policy probabilities, but estimating the hindsight distribution requires an auxiliary model or an extra pass. To address this estimation bottleneck, we propose GraphHCA, a model-free realization of HCA that eliminates explicit hindsight-distribution estimation. For terminal-goal tasks with deterministic transitions, Bayes' rule reduces the hindsight ratio to a ratio of behavior-policy success probabilities at consecutive states. Taking logs yields a state-wise success potential, whose increment across a transition provides step-level credit. GraphHCA estimates this potential from pooled rollouts through a discounted recursion on the induced transition graph, which admits a unique fixed point on any directed graph. The resulting step-level signal is combined with the trajectory-level advantage, requiring neither a learned hindsight model nor an extra forward pass and recovering GRPO when the step-level weight is zero. Among all compared baselines, GraphHCA achieves state-of-the-art results on ALFWorld and WebShop at both LLM scales, and on Sokoban with a vision-language agent. For example, on ALFWorld it improves overall success rate by up to 24.6 points over GRPO and by up to 4.7 points over the strongest step-level baseline.
Haodong Zhu, Yangyang Ren, Changbai Li +4
Beihang University · Zhongguancun Academy · Communication University of China +1
Reinforcement learning has become a powerful paradigm for post-training large language model agents, yet credit assignment in multi-turn environments remains a challenge. Agents often receive sparse, trajectory-level rewards only at the end of an episode, making it difficult to determine which intermediate actions contributed to success or failure. As a result, propagating delayed outcomes back to individual decision steps without relying on costly auxiliary value models remains an open problem. We propose Generalized Advantage Grouped Policy Optimization (GAGPO), a critic-free reinforcement learning method for precise, step-aligned temporal credit assignment. GAGPO constructs a non-parametric grouped value proxy from sampled rollouts and uses it to compute TD/GAE-style temporal advantages, recursively propagating outcome supervision backward through time. Combined with group-wise advantage normalization and an action-level importance ratio, GAGPO extracts stable, localized optimization signals directly from multi-turn trajectories. Experiments on ALFWorld and WebShop show that GAGPO outperforms strong reinforcement learning baselines. Further analyses demonstrate faster early-stage learning, improved interaction efficiency, and smoother optimization dynamics, suggesting that GAGPO offers a simple yet effective framework for multi-turn agentic reinforcement learning.
Siyuan Zhu, Chao Yu, Rongxin Yang +4
School of Computer Science and Engineering, Sun Yat-sen University · Meituan