cs.AISep 24, 2026

SLCA-GRPO: Resolving Cross-Segment Credit Misattribution in Tool-Calling RL

Authors: Yan Zhan, Shaobo Liu, Qiunan Liu, Yuanjun Shi, Siqi Xu, WeiYi Hou, Xiang Xu, Zekang Li, +2 more

Organizations: Peking University · Shenzhen University · Tencent PCG QQ Team

Abstract

Tool-calling agents produce heterogeneous outputs, interleaving structured tool invocations with user-facing natural language summaries. This output heterogeneity presents a structural failure mode in standard on-policy Reinforcement Learning (RL): algorithms like GRPO indiscriminately broadcast a homogeneous trajectory-level scalar advantage to all tokens. Consequently, gradient noise from summary generation leaks into tool-decision tokens, causing cross-segment credit misattribution and brittle optimization. In this work, we propose SLCA-GRPO, a framework incorporating Segment-Locked Credit Assignment (SLCA). To enable scalable exploration without costly real APIs and stable training, we first construct the Schema-Guided LLM Simulator (SGLS) as foundational training infrastructure. Building on this, SLCA decouples advantage estimation at the structural segment level within a single group of rollouts, without requiring additional rollouts from intermediate states. Supported by Hierarchical Rewards (HierR), SLCA routes execution advantages to tool tokens and preference advantages to summary tokens, eliminating advantage contamination (the dominant cross-segment credit misattribution channel) within each policy update. On a 7B backbone, SLCA-GRPO accelerates convergence and outperforms standard GRPO, ToolPO, and RLTR by +2.53 pp on in-domain evaluation, +1.36 pp on the Berkeley Function-Calling Leaderboard (BFCL), and +9.15 pp on τ2τ^2-Bench under the same training budgets, achieving higher accuracy with reduced tool redundancy and costs.

Figures & tables

Appendix figures & tables30 assets

Supplementary material from the paper’s appendix.

Appendix

Explore similar work

CardsList
  1. SHARPO: Segment-Level Credit Assignment for Agentic Reinforcement Learning

    Sep 30, 2026Xinchen Du, Zhengze Zhou, Wenhui Zhu +4Agentic Reinforcement LearningCredit Assignment

  2. APPO: Agentic Procedural Policy Optimization

    Jun 10, 2026Xucong Wang, Ziyu Ma, Yong Wang +5Agentic Reinforcement LearningFrictive Policy Optimization

  3. SFT or RL for Tool-Calling Agents? A Controlled Study Across Data, Method, and Scale

    Sep 15, 2026Md Tahmid Rahman Laskar, Xue-Yong Fu, Shashi Bhushan TNLanguage-Model AgentsFlow-Grpo