Tool-calling agents produce heterogeneous outputs, interleaving structured tool invocations with user-facing natural language summaries. This output heterogeneity presents a structural failure mode in standard on-policy Reinforcement Learning (RL): algorithms like GRPO indiscriminately broadcast a homogeneous trajectory-level scalar advantage to all tokens. Consequently, gradient noise from summary generation leaks into tool-decision tokens, causing cross-segment credit misattribution and brittle optimization. In this work, we propose SLCA-GRPO, a framework incorporating Segment-Locked Credit Assignment (SLCA). To enable scalable exploration without costly real APIs and stable training, we first construct the Schema-Guided LLM Simulator (SGLS) as foundational training infrastructure. Building on this, SLCA decouples advantage estimation at the structural segment level within a single group of rollouts, without requiring additional rollouts from intermediate states. Supported by Hierarchical Rewards (HierR), SLCA routes execution advantages to tool tokens and preference advantages to summary tokens, eliminating advantage contamination (the dominant cross-segment credit misattribution channel) within each policy update. On a 7B backbone, SLCA-GRPO accelerates convergence and outperforms standard GRPO, ToolPO, and RLTR by +2.53 pp on in-domain evaluation, +1.36 pp on the Berkeley Function-Calling Leaderboard (BFCL), and +9.15 pp on τ2-Bench under the same training budgets, achieving higher accuracy with reduced tool redundancy and costs.
Figures & tables
Diagnostic cases. Unified reward aggregation can misassign credit across tool and summary segments. Left: an unnecessary tool call is paired with a correct delivery summary. Right: correct tool evidence is paired with an incorrect refund summary. Under standard GRPO, the mixed advantage can reward bad tool use or penalize good tool use.
Figure 1 : Overview of the SLCA -GRPO Framework. (Top) Infrastructure: SGLS enables scalable exploration via schema-constrained simulation, generating interleaved trajectories. (Middle) Signal: HierR decouples feedback into dense execution rewards ( {\color[rgb]{0.418,0.3711,0.8203}R^{\mathrm{tool}}} ) and terminal outcome rewards ( {\color[rgb]{0.3711,0.6406,0.3047}R^{\mathrm{sum}}} ). (Bottom) Optimization: SLCA decouples segment-wise advantages. By normalizing and routing advantages independently ( {\color[rgb]{0.418,0.3711,0.8203}\hat{A}^{\mathrm{tool}}} vs. {\color[rgb]{0.3711,0.6406,0.3047}\hat{A}^{\mathrm{sum}}} ), it blocks the defined summary-to-tool support path within each policy update.
Method
Name F1
ArgMatch
Process
Success
Qwen2.5-3B-Instruct
Original
0.5207
0.5073
0.5174
0.3816
SFT
0.7685 ± .0069
0.7132 ± .0239
0.7401 ± .0050
0.6500 ± .0089
SFT+GRPO
0.8910 ± .0087
0.8072 ± .0154
0.8577 ± .0057
0.7412 ± .0115
RLTR
0.7598 ± .0149
0.6353 ± .0229
0.6924 ± .0162
0.6627 ± .0151
ToolPO
0.8050 ± .0142
0.7769 ± .0160
0.8368 ± .0141
0.5128 ± .0145
Table 1 : Toucan-Test Results. Metrics: Name F1: Multiset F1 score measuring precision and recall of predicted tool-name occurrences against gold names; ArgMatch: Arithmetic mean of argument key and value matching scores; Process: The dense tool-segment reward ( Sprocess ) defined in Eq. 22 (App. A.4 ); Success: Strict binary indicator ( I[Sprocess≥0.9] ) obtained by thresholding the process score. All trained entries report mean ± std over three runs; Original rows are point evaluations. The matched rows share the per-backbone SFT initialization, data split, SGLS endpoint, mocker configuration, decoding settings, evaluation protocol, G=16 , one RL epoch, and run set. GRPO, SLCA -GRPO , w/o SLCA, and w/o SGLS also share the HierR definition and subweights; w/o HierR removes HierR. ToolPO and RLTR retain their method-specific protocols. The ToolPO rows use the LLM-judge outcome reward; App. C.4 reports how this baseline responds to a rule-based outcome reward under the same estimator.
Figure 2 : Training Dynamics. (a) Baseline comparison across 3B/7B/8B for SLCA / ToolPO / RLTR (the ToolPO diagnostic is described in App. D.3 ), and (b) cross-scale dynamics of success rate and average tool turns over training steps.
Figure 3 : OOD Generalization across Scales. (a) Atomic generalization on BFCL V3 and (b) Collaboration robustness on τ2 -Bench for the three backbones (3B, 7B, 8B), comparing SLCA -GRPO against standard GRPO, RLTR, and ToolPO. Per-domain τ2 -Bench details are in App. D.6 .
Toucan-Test
BFCL
τ2 -Bench
Setting
Process
Success
Acc
Pass 1
SLCA -GRPO
0.8766 ± .0045
0.7913 ± .0105
0.6977 ± .0047
0.4102 ± .0101
w/o SLCA
0.8667 ± .0061
0.7660 ± .0127
0.6841 ± .0011
0.3187 ± .0297
w/o SGLS
0.8787 ± .0051
0.7802 ± .0112
0.6875 ± .0048
0.3619 ± .0195
w/o HierR
0.8521 ± .0083
0.7344 ± .0121
0.6907 ± .0046
0.3450 ± .0144
Table 2: Ablation Studies on Qwen2.5-7B-Instruct. Entries report mean ± std over three independent runs. Full results are provided in App. D.10 . Process is an evaluation metric computed post hoc for all settings, including w/o HierR.
Appendix figures & tables30 assets
Supplementary material from the paper’s appendix.
Appendix
Figure 4 : Mechanism of Cross-Segment Credit Misattribution and segment-locked routing. Unified advantage broadcast leaks summary rewards into tool tokens, whereas SLCA blocks this path by routing segment-wise advantages only to their corresponding tokens.
Parameter
Value
Model Specification
Model
GPT-OSS-120B
Parameters
120B
Serving Framework
vLLM (OpenAI-compatible API)
Decoding
Temperature
0.0 (fully greedy / deterministic)
Appendix
Table 3: Summary Judge: Deployment and Decoding Parameters.
Figure 5 : Gradient Norm Dynamics ( ∥∇θ∥2 ) during Training. We compare the gradient norms of SFT+GRPO (w/o SLCA) (Orange) and SLCA -GRPO (Blue) across Qwen2.5-3B, 7B, and Qwen3-8B. These are representative single-run traces and provide a descriptive view of update variability under the two estimators.
Model Backbone
Ours Std ( σSLCA )
Baseline Std ( σBase )
Improvement ( ↓ )
Qwen2.5-3B-Instruct
0.156
0.209
25.0%
Qwen2.5-7B-Instruct
0.094
0.108
13.1%
Qwen3-8B-Base
0.312
0.320
2.5%
Appendix
Table 5: Quantification of Gradient Stability. We report the Standard Deviation (Std) of the gradient norm. SLCA -GRPO has lower observed volatility in these representative runs.
Figure 6 : Data Processing Pipeline. The workflow applies strict relevance filtering and format check before isolating the evaluation set. The remaining training pool is split into SFT and RL partitions; the RL set combines native single-turn and decomposed multi-turn trajectories formatted for the SGLS.
Fine-Tuning Hyperparameters
RL Hyperparameters
Parameter
Value
Parameter
Value
Hardware Resources
8 × H20 GPUs
Hardware Resources
32 × H20 (4 Nodes)
Optimizer
AdamW
Precision
bfloat16
Total Batch Size
32
Total Batch Size
128
Per-Device Batch
1
Mini-batch Size
32
Grad Accumulation
4
Rollout Group Size ( G )
16
Appendix
Table 6: Detailed Hyperparameters and Environment Settings. We list the configurations for both the supervised fine-tuning (SFT) and the subsequent RL training stages.
Parameter
Value
vLLM Deployment
Tensor-Parallel Size (TP)
8
Max Model Length
32,768
GPU Memory Utilization
0.86
Max Concurrent Sequences
300
Precision (dtype)
bfloat16
Appendix
Table 7: SGLS Frozen LLM: Deployment and Decoding Parameters.
Figure 7 : SGLS mock responses align with real-API outputs across three representative tool domains. The Semantic Mapping ϕ(o) extracts schemas, facts, and logic, realizing three levels of alignment (Mapping Marks): Identical Value for deterministic tools ( calculator ), Identical Schema for query-style tools ( weather ; individual values need not match a specific real-API snapshot but remain within plausible ranges), and Semantic Equivalence for open-ended tools ( search ; factually consistent content differing only in phrasing).
Method
Toucan Succ. (%)
BFCL Acc. (%)
τ2 -Bench Pass 1 (%)
SFT+GRPO (w/o SLCA)
75.74 ± 1.47
68.08 ± 0.36
32.74 ± 2.66
SLCA -GRPO
78.57 ± 1.15
69.32 ± 0.56
39.90 ± 1.46
Δ
+2.83
+1.24
+7.16
Appendix
Table 8: Response-mocker check on Qwen2.5-7B-Instruct. GPT-OSS-120B replaces the response mocker; entries are mean ± standard deviation under the matched policy setup.
Condition
Toucan Succ. (%)
BFCL Acc. (%)
τ2 -Bench Pass 1 (%)
ToolPO, LLM-judge Rsuccjudge
20.00 ± 1.36
24.59 ± 0.87
30.44 ± 2.15
ToolPO, reference rule Rsuccrule
77.35 ± 1.19
68.86 ± 0.43
34.18 ± 2.42
Unified GRPO
76.60 ± 1.27
68.41 ± 0.11
31.87 ± 2.97
SLCA -GRPO
79.13 ± 1.05
69.77 ± 0.47
41.02 ± 1.01
Appendix
Table 9: ToolPO outcome reward protocol comparison on Qwen2.5-7B-Instruct. Entries are mean ± standard deviation over three matched runs. The reference rule row uses Eq. 28 ; Unified GRPO and SLCA -GRPO are shown as contextual matched baselines under their standard reward protocol.
Parameter
Value
tau2 package
v0.2.1.dev0 (from @v0.2.0 tag)
Evaluation framework
evalscope 1.3.0
Task split
airline, retail, telecom (all 3 official domains)
Dataset source
evalscope/tau2-bench-data (ModelScope)
Task filtering
LLMGTAgent.check_valid_task()
Aggregation
mean_and_pass_hat_k (built-in)
Appendix
Table 10: τ2 -Bench evaluation protocol.
Parameter
Value
Evaluation framework
evalscope 1.3.0
Evaluation mode
is_fc_model=True
Temperature
0 (greedy)
Max tokens
8192
parallel_tool_calls
True
underscore_to_dot
True
Appendix
Table 11: BFCL V3 evaluation protocol.
Parameter
Value
Dataset
tool_call_rl_40k_v1.parquet (4,000 held-out)
Agent loop max turns
10
Temperature
0.0 (greedy)
Max tokens
12,288
Top- p
1.0
Concurrency
256
Appendix
Table 12: Toucan in-domain evaluation protocol.
Candidate Size ( K )
Original
20
50
100
150
Valid Samples
3,991
3,980
3,882
2,607
443
Avg. Tokens
943
4,614
11,064
17,177
23,635
Std. Dev.
1,114
3,679
5,256
1,550
733
Skip Rate
0.2%
0.5%
2.9%
34.8%
88.9%
Appendix
Table 13: System prompt length statistics vs. candidate set size ( K ) for Random distractors (Qwen2.5 tokenizer).
Candidate Size ( K )
Original
20
50
100
150
Valid Samples
3,991
3,959
3,793
3,023
2,034
Avg. Tokens
943
3,693
8,522
14,483
20,609
Std. Dev.
1,114
2,735
3,806
2,595
1,719
Skip Rate
0.2%
1.0%
5.2%
24.4%
49.2%
Appendix
Table 14: System prompt length statistics vs. candidate set size ( K ) for Hard distractors (Qwen2.5 tokenizer).
Method
Name F1
ArgMatch
Parallel
Process
Summary
Success
Qwen2.5-3B-Instruct
Original
0.5207
0.5073
0.5412
0.5174
0.3477
0.3816
SFT
0.7685 ± .0069
0.7132 ± .0239
0.7772 ± .0071
0.7401 ± .0050
0.5512 ± .0047
0.6500 ± .0089
SFT+GRPO
0.8910 ± .0087
0.8072 ± .0154
0.9176 ± .0080
0.8577 ± .0057
0.7842 ± .0122
0.7412 ± .0115
RLTR
0.7598 ± .0149
0.6353 ± .0229
0.6729 ± .0098
0.6924 ± .0162
–
0.6627 ± .0151
ToolPO
0.8050 ± .0142
0.7769 ± .0160
0.8946 ± .0137
0.8368 ± .0141
0.6189 ± .0117
0.5128 ± .0145
Appendix
Table 15: Comprehensive Results on Toucan-Test across Scales. This table expands upon the main results by including all sub-metrics. In addition to Name F1, ArgMatch, Process, and Success (defined in Table 1 ), we report: Parallel: The parallelism/cardinality constraint score defined in Eq. 27 ; Summary: The user-facing response quality score evaluated by the LLM-judge. SLCA -GRPO has the highest mean Success and Process values across scales. RLTR and ToolPO are included for context; their method-specific protocols are described in App. C.4 . RLTR’s Summary column is marked “–” as its frozen summarizer is not optimized during RL. Trained rows report mean ± std over three runs; Original rows are point evaluations.
Condition
No-call (%)
Process
Toucan (%)
Summary
BFCL (%)
τ2 (%)
SFT
–
0.8094 ± 0.0082
72.14 ± 0.87
0.6346 ± 0.0054
67.89 ± 0.28
32.54 ± 1.92
Unified GRPO
0.6
0.8667 ± 0.0061
76.60 ± 1.27
0.8241 ± 0.0108
68.41 ± 0.11
31.87 ± 2.97
A^sum=0 (except omission penalty episodes)
0.5
0.8748 ± 0.0061
78.72 ± 1.22
0.6512 ± 0.0208
69.44 ± 0.55
36.24 ± 1.58
A^sum=0 (including omission penalty episodes)
14.2
0.7506 ± 0.0227
67.54 ± 1.94
0.5587 ± 0.0248
61.53 ± 1.74
29.44 ± 2.28
SLCA -GRPO
0.4
0.8766 ± 0.0045
79.13 ± 1.05
0.8450 ± 0.0087
69.77 ± 0.47
41.02 ± 1.01
RLTR
–
0.7022 ± 0.0143
65.45 ± 1.68
–
63.11 ± 0.62
33.72 ± 2.38
Appendix
Table 16 : Summary advantage ablation on Qwen2.5-7B-Instruct. Process and Summary are unit-interval scores; benchmark columns are percentages. The no-call column is the end-of-training rate. Entries in the score columns are mean ± standard deviation over three matched runs.
Figure 8 : ToolPO training behavior under the LLM-judge diagnostic (representative single-run traces). (a) Success@0.9 across three scales: a decline on 3B ( 0.65→0.51 ), a sharper decline on 7B ( 0.72→0.20 ), and an overall decline with a brief mid-training rebound on 8B, ending near 0.32 . (b) 7B: tag_open_n diverges from tag_close_n after step 80, reducing the valid parsed-call count from 1.91 to 0.73. (c) 8B: pred_tool_count surges 1.71→3.98 (gold ≈2.0 ), driving Name F1 and Parallel Score down. (d) 3B: Process Score plateaus at ∼0.84 , below the Sprocess≥0.9 threshold, explaining why proxy metrics improve while strict Success declines.
Figure 9 : Diagnostic Case Study: Resolving Cross-Segment Credit Misattribution on a parallel tool-use task. (Top) Reference trajectory and rewards. (Middle) Standard GRPO. The agent adopts a redundant, trial-and-error trajectory (first issuing an unnecessary search_company_info call, then retrieving the two stock prices through serial get_stock_price invocations in separate turns), yet still produces a fluent, factually correct comparison ( {\color[rgb]{0.3711,0.6406,0.3047}R_{\mathrm{sum}}}{=}1.00 ). The tool-level decomposition exposes where the trajectory fails: the format reward is unaffected ( 1.00 ), the key and value rewards remain high on the individual calls that do match the schema ( 1.00 each), the name reward is diluted by the redundant search_company_info call ( 0.80 ), and the parallel reward is 0.00 because the two required stock-price queries are issued as separate sequential turns (further preceded by a redundant exploration call), rather than batched into a single parallel invocation. Aggregated under the fixed weights above, this yields {\color[rgb]{0.418,0.3711,0.8203}R_{\mathrm{tool}}}{=}0.65 . Because standard GRPO merges {\color[rgb]{0.418,0.3711,0.8203}R_{\mathrm{tool}}} and {\color[rgb]{0.3711,0.6406,0.3047}R_{\mathrm{sum}}} into a single trajectory-level advantage, the high summary reward can offset the structural tool penalty in the unified advantage, allowing the inefficient pattern to be reinforced. (Bottom) SLCA -GRPO (Ours). The same model, trained under our advantage-isolation objective, produces a trajectory in which both tickers are known a priori and issues a single parallel tool call covering AAPL and MSFT simultaneously. All five tool sub-rewards saturate at 1.00 , giving {\color[rgb]{0.418,0.3711,0.8203}R_{\mathrm{tool}}}{=}1.00 while preserving {\color[rgb]{0.3711,0.6406,0.3047}R_{\mathrm{sum}}}{=}1.00 . By routing {\color[rgb]{0.418,0.3711,0.8203}A_{\mathrm{tool}}} and {\color[rgb]{0.3711,0.6406,0.3047}A_{\mathrm{sum}}} through separate optimization pathways, segment routing keeps the summary signal from entering the tool-token update in this example. The tool-execution score increases by +0.35 while the summary score remains 1.00 . The trajectory is qualitatively similar to the representative single-run pred_tool_count increase observed under ToolPO on 8B ( Figure 8 c); the two diagnostics use different protocols.
Method
Overall Acc
AST Live
AST Non-Live
Multi-Turn
Qwen2.5-3B-Instruct
Original
0.4225
0.5174
0.5139
0.0271
SFT
0.6163 ± .0026
0.6862 ± .0022
0.8104 ± .0043
0.0867 ± .0100
SFT+GRPO
0.6321 ± .0032
0.6936 ± .0030
0.8583 ± .0087
0.0600 ± .0033
RLTR
0.5838 ± .0065
0.6136 ± .0074
0.8275 ± .0044
0.0494 ± .0100
ToolPO
0.6095 ± .0086
0.6477 ± .0108
0.8325 ± .0139
0.0961 ± .0067
Appendix
Table 17: Comprehensive Generalization Results on BFCL V3 across Scales. This table details BFCL generalization performance. In addition to Overall Acc, AST Live, and AST Non-Live (reported in Figure 3(a) ), we include Multi-Turn accuracy to assess state tracking under schema generalization. Trained rows report mean ± std over three runs; Original rows are point evaluations. ToolPO and RLTR retain their method-specific protocols.
Method
Overall Pass 1
Airline
Retail
Telecom
Qwen2.5-3B-Instruct
Original
0.2324
0.3810
0.1481
0.2692
SFT
0.2705 ± .0236
0.1364 ± .0227
0.3966 ± .0259
0.1954 ± .0217
SFT+GRPO
0.3454 ± .0257
0.2576 ± .0131
0.4598 ± .0217
0.2644 ± .0348
RLTR
0.2850 ± .0291
0.2727 ± .0227
0.3649 ± .0303
0.2098 ± .0303
ToolPO
0.2355 ± .0226
0.1288 ± .0262
0.3678 ± .0263
0.1437 ± .0179
Appendix
Table 18: Comprehensive Collaboration Results on τ2 -Bench across Scales. We report Pass 1 rates for dual-control scenarios across three distinct domains: Airline (User Retention), Retail (Order Modification), and Telecom (Plan Management). Pass 1 is the proportion of successful tasks in the finite official task set; because each run evaluates a finite set, the reported values are discrete proportions and their standard deviations reflect variation of those proportions across runs. The Qwen3-8B-Base Original row is included and is zero in all four reported metrics. Trained rows are mean ± std over three runs. ToolPO and RLTR retain their method-specific protocols.
Sprocess Weights
Config
Sfmt
Sname
Skey
Sval
Spar
BFCL Acc (%)
τ2 Pass 1 (%)
Default (SLCA)
0.10
0.25
0.15
0.20
0.30
69.77 ± 0.47
41.02 ± 1.01
Uniform (SLCA)
0.20
0.20
0.20
0.20
0.20
69.18 ± 0.55
38.62 ± 1.86
Value-heavy (SLCA)
0.05
0.20
0.20
0.35
0.20
69.44 ± 0.48
39.35 ± 1.73
w/o HierR
—
69.07 ± 0.46
34.50 ± 1.44
GRPO (unified adv.)
same as Default
68.41 ± 0.11
31.87 ± 2.97
Appendix
Table 19: HierR weight sensitivity on Qwen2.5-7B-Instruct. Entries are mean ± std over three runs.
Reward ratio
Toucan Success@0.9 (%)
BFCL Acc. (%)
τ2 -Bench Pass 1 (%)
Unified 1:1
76.60 ± 1.27
68.41 ± 0.11
31.87 ± 2.97
Unified 2:1
77.54 ± 1.24
68.83 ± 0.31
34.62 ± 2.71
Unified 3:1
77.88 ± 1.30
68.95 ± 0.36
35.24 ± 2.58
Unified 5:1
77.02 ± 1.41
68.60 ± 0.44
32.90 ± 2.88
SLCA -GRPO
79.13 ± 1.05
69.77 ± 0.47
41.02 ± 1.01
Appendix
Table 20: Unified reward ratio sensitivity on Qwen2.5-7B-Instruct. Entries are mean ± standard deviation over three matched runs. The ratio is applied before the unified group normalization; the scale-invariance identity holds up to the numerical floor described in App. A.3 .
Condition
Tool tokens
Summary tokens
Toucan (%)
BFCL (%)
τ2 (%)
Both closed (SLCA)
T
S
79.13 ± 1.05
69.77 ± 0.47
41.02 ± 1.01
S→T open
T+S
S
77.21 ± 1.24
68.72 ± 0.29
33.95 ± 2.71
T→S open
T
T+S
79.02 ± 1.12
69.84 ± 0.44
41.36 ± 1.24
Both open
T+S
T+S
78.30 ± 1.09
69.30 ± 0.41
37.42 ± 2.10
Unified (joint normalization)
joint normalized reward
joint normalized reward
76.60 ± 1.27
68.41 ± 0.11
31.87 ± 2.97
Appendix
Table 21 : Support control on Qwen2.5-7B-Instruct. The four support cells use separately normalized components. Unified is a joint-normalization reference. Entries are mean ± standard deviation over three matched runs.
Effect
Toucan
BFCL
τ2 -Bench
Close S→T
1.32
0.80
5.51
Open T→S
0.49
0.33
1.91
Interaction
1.20
0.51
3.13
Appendix
Table 22: Factor effects from the 7B support control. Values are computed from the cell means in Table 21 and are reported in percentage points. Main effects average the two simple effects over the other factor; interaction entries use the unhalved difference-in-differences contrast.
Method
Toucan-Test
BFCL
τ2 -Bench
Name F1
ArgMatch
Process
Success
Acc
Pass 1
Qwen2.5-3B-Instruct
SLCA -GRPO
0.9006 ± .0087
0.8092 ± .0112
0.8625 ± .0049
0.7647 ± .0093
0.6361 ± .0052
0.3563 ± .0091
w/o SLCA
0.8910 ± .0087
0.8072 ± .0154
0.8577 ± .0057
0.7412 ± .0115
0.6321 ± .0032
0.3454 ± .0257
w/o SGLS
0.8927 ± .0075
0.7982 ± .0109
0.8559 ± .0050
0.7584 ± .0101
0.6304 ± .0048
0.2995 ± .0182
w/o HierR
0.8399 ± .0084
0.7850 ± .0161
0.8264 ± .0075
0.6552 ± .0107
0.5993 ± .0046
0.3152 ± .0145
Appendix
Table 23: Comprehensive Ablation Studies Across Model Families. Comparison of Full SLCA -GRPO against variants removing key components. Metrics include Toucan-Test (Name F1, ArgMatch, Process, Success), BFCL (Accuracy), and τ2 -Bench (Pass 1 ). Process is computed as a post-hoc evaluation metric for all settings, including w/o HierR. Trained rows in all three blocks report mean ± std over three runs; Original evaluations are point values where shown. The w/o SLCA condition is the SFT+GRPO estimator.
Candidate Tool Set Size ( K )
Method
Original
20
50
100
150
Hard-Negative Sampling Qwen2.5-3B-Instruct
SFT+GRPO
0.7412
0.3400
0.2921
0.2662
0.2285
Ours
0.7647
0.3584
0.3209
0.2865
0.2510
Random Sampling Qwen2.5-3B-Instruct
SFT+GRPO
0.7412
0.6399
0.5701
0.4950
0.4360
Appendix
Table 24: Comprehensive Scalability Results Across Model Families. We report Success @ 0.9 as the candidate set size K increases. Comparison is made between Hard-Negative (Semantic) and Random sampling to isolate the impact of semantic interference. The Original column reports the multi-run mean at K=0 , matching the main results; the K>0 columns retain the stress-test evaluations. For K≥100 , especially K=150 , these stress-test values are computed on context-valid survivors after skip filtering.
Method
Toucan Succ. (%)
BFCL Acc. (%)
τ2 -Bench Pass 1 (%)
SFT+GRPO (w/o SLCA)
73.51 ± 1.45
66.28 ± 0.32
25.54 ± 2.13
SLCA -GRPO
76.43 ± 1.21
67.61 ± 0.51
30.18 ± 1.76
Δ
+2.92
+1.33
+4.64
Appendix
Table 25: Execution-based Rsucc comparison on Qwen2.5-7B-Instruct. Entries are mean ± standard deviation over three runs; data, initialization, SGLS, and compute are matched.
Figure 10 : Training dynamics under execution-based Rsucc (7B). The curves are single-run checkpoint evaluations on the same fixed set of 3,991 valid Toucan validation examples. The plotted Rsucc is distinct from Toucan Success@0.9. At step 100, SLCA -GRPO reaches Rsucc=0.593 , compared with 0.424 for SFT+GRPO (w/o SLCA).
Agentic reinforcement learning (RL) trains a large language model (LLM) to act over long, multi-step interactions. However, a single localized error can cause task failure, while trajectory-level rewards provide limited guidance for assigning credit to individual decisions. To address this limitation, we introduce Segment-level Hindsight Advantage Reweighting for Policy Optimization (SHARPO), a credit-assignment mechanism that refines Group Relative Policy Optimization (GRPO) at the level of environment-facing segments. Inspired by the existing on-policy self-distillation (OPSD) method, SHARPO computes teacher-student log-probability gaps within each segment and uses the resulting signal to compute a bounded multiplier on the GRPO advantage. This multiplier is shared by all tokens within the segment, allowing credit to vary across different segments. With Qwen2.5-7B-Instruct, SHARPO outperforms existing baselines on the ALFWorld and WebShop benchmarks, including GRPO, SDAR, RLSD, and StepOPSD.
Xinchen Du, Zhengze Zhou, Wenhui Zhu +4
LinkedIn Corporation · Georgia Institute of Technology
Recent advances in agentic Reinforcement Learning (RL) have substantially improved the multi-turn tool-use capabilities of large language model agents. However, most existing methods assign credit over coarse heuristic units, such as tool-call boundaries or fixed workflows, making it difficult to identify which intermediate decisions influence downstream outcomes. In this work, we study agentic RL from two perspectives: \textit{where to branch and how to assign credit after branching}. Our pilot analysis shows that influential decision points are broadly distributed throughout the generated sequence rather than concentrated at tool calls, while token entropy alone does not reliably reflect their impact on final outcomes. Motivated by these observations, we propose \textbf{Agentic Procedural Policy Optimization (APPO)}, which shifts branching and credit assignment from coarse interaction units to fine-grained decision points in the sequence. APPO selects branching locations using a Branching Score that combines token uncertainty with policy-induced likelihood gains of subsequent continuations, enabling more targeted exploration while filtering out spurious high-entropy positions. It further introduces procedure-level advantage scaling to better distribute credit across branched rollouts. Experiments on 13 benchmarks show that APPO consistently improves strong agentic RL baselines by nearly 4 points, while keeping efficient tool-calls and maintaining behavior interpretability.
Xucong Wang, Ziyu Ma, Yong Wang +5
University of Science and Technology of China · AMAP, Alibaba Group · Southern University of Science and Technology
Limited controlled evidence exists on how training data, adaptation method, and model scale jointly affect tool-calling performance in language-model agents. We evaluate supervised fine-tuning (SFT) with LoRA, reinforcement learning (RL) via Group Relative Policy Optimization (GRPO), and SFT followed by GRPO across six Qwen3 models from 0.6B to 32B parameters, covering both in-distribution performance and cross-dataset transfer. SFT with LoRA is the strongest in-distribution method throughout the 0.6B-32B range and best in 15 out of 18 experimental settings. On cross-dataset transfer, the methods are closer: GRPO wins 29 out of 54 settings where training and test datasets differ, but its margin over SFT averages under one point, and SFT->GRPO is rarely strongest in either comparison. Dataset mixing gives consistently strong transfer while staying close to specialized in-distribution training, regardless of method. Additional analysis further confirms that LoRA outperforms full-parameter fine-tuning, demonstrating that LoRA better preserves pretrained agentic behavior.