cs.LGOct 5, 2026

RELACE: retrospective likelihood-based action credit estimation for long-horizon language agents

Authors: Sayak Chakrabarti, Sathish Reddy Indurthi

Organizations: Columbia University · Zoom Communications, Inc.

Abstract

Group Relative Policy Optimization (GRPO) avoids a separate critic by estimating advantages from rollout groups. For multi-turn agents, however, trajectory-level supervision provides coarse, noisy credit: terminal rewards do not locate errors and can penalize useful actions alongside mistakes. Group-in-Group Policy Optimization (GiGPO) and subsequent methods refine supervision through state-conditioned comparisons, but their credit estimates remain sensitive to downstream decisions and outcomes. We introduce RELACE, Retrospective Likelihood-based Action, a critic-free framework that integrates retrospective action assessment with state-conditioned advantage estimation. RELACE evaluates executed actions through teacher-forced likelihood scoring under both their original contexts and outcome-augmented contexts. Comparing these likelihoods yields a trajectory-normalized retrospective factor that captures outcome-dependent changes in action plausibility, rather than hindsight plausibility alone. We use this factor to reweight discounted task returns and construct local advantages by comparing weighted returns among actions from equivalent states within a task. This couples retrospective relevance with observed reward, producing fine-grained credit that complements trajectory-level GRPO supervision. Temporal smoothing and success-protecting masking further stabilize the local signal. RELACE requires neither auxiliary value nor reward models nor additional autoregressive rollouts for credit estimation. Experiments on ALFWorld and WebShop with Qwen2.5-1.5B-Instruct and Qwen2.5-7B-Instruct demonstrate substantial improvements over GRPO, GiGPO, and HCAPO. With the 1.5B model, RELACE achieves 96.35%96.35\% success on ALFWorld and 79.43%79.43\% on WebShop, surpassing GiGPO by 5.475.47 and 5.605.60 percentage points, respectively.

Figures & tables

Appendix figures & tables1 asset

Supplementary material from the paper’s appendix.

Appendix

Explore similar work

CardsList
  1. Granularity-Adaptive Credit Assignment for Long-Horizon LLM Agent Reinforcement Learning

    Sep 11, 2026Taoran Liang, Yang Liu, Shang Luo +10Credit Assignment in RLStep-Level Credit Assignment in RL

  2. GraphHCA: Closed-Form Hindsight Credit Assignment for Long-Horizon LLM Agents

    Sep 28, 2026Haodong Zhu, Yangyang Ren, Changbai Li +4Credit Assignment in RLStep-Level Credit Assignment in RL

  3. GAGPO: Generalized Advantage Grouped Policy Optimization

    May 13, 2026Siyuan Zhu, Chao Yu, Rongxin Yang +4Reinforcement LearningGroup Relative Policy Optimization