Multi-turn LLM agents often receive sparse task feedback across several interactions, while generating each response token by token. This creates two related credit-assignment questions: which responses helped achieve the outcome, and which generation decisions mattered within each response? Existing methods typically focus on only one level: turn-level methods evaluate complete responses but do not distinguish the decisions within them; token-level methods can propagate feedback across turns but do not explicitly model credit for each response. These complementary limitations motivate learning credit at both levels and coordinating it in a single policy update. We introduce VETTA, a credit assignment method that jointly learns turn- and token-level values through separate heads on a shared lightweight critic. VETTA computes advantages along both temporal sequences and combines each turn advantage with a within-response-centered token residual for PPO updates. Furthermore, to reduce value-learning cost, the critic retains only early Transformer blocks from the pretrained checkpoint used to initialize the actor. On two challenging agent benchmarks, ALFWorld and WebShop, VETTA improves success rates over PPO by 37.5% and 22.3%, respectively, with Qwen2.5-1.5B-Instruct and achieves success rates of 95.5% and 76.0%, respectively, with Qwen2.5-7B-Instruct. Critic-depth comparisons further show strong task performance with substantially lower critic-side computation. These results suggest that a compact shared critic can coordinate turn- and token-level credit to improve agent performance while keeping value estimation efficient. Code is available at https://github.com/Jiaju-Chen/VETTA-official.
Figures & tables
Figure 1: Success rates on ALFWorld and WebShop: (a) full-depth PPO*, Turn-PPO, and mixed configurations; (b) the mixed configuration with 2, 8, or 28 critic blocks. Bars show means and error bars show sample standard deviations over three decoding seeds for each checkpoint.
Figure 2: Overview of VETTA. (a) A shared critic retains early pretrained Transformer blocks and uses separate heads for turn- and token-level value estimation. (b) The actor interacts with the environment to collect trajectories and rewards; value estimates support advantage computation, followed by critic regression and PPO policy updates. (c) The collected trajectory is viewed at turn and token levels for GAE; token advantages are centered within each response and combined with its turn advantage. The implicit successor value and advantage are set to zero at rollout end.
Type
Method
ALFWorld
WebShop
Pick
Look
Clean
Heat
Cool
Pick2
All
Score
Succ.
Qwen2.5-1.5B-Instruct
Prompting
Qwen2.5 †
5.9
5.5
3.3
9.7
4.2
0.0
4.1
23.1
5.2
Prompting
ReAct †
17.4
20.5
15.7
6.2
7.7
2.0
12.8
40.1
11.3
Prompting
Reflexion †
35.3
22.2
21.7
13.6
19.4
3.7
21.8
55.8
21.9
Critic-free
RLOO †
88.3 ±3.0
52.8 ±8.6
71.0 ±5.9
62.8 ±8.7
66.4 ±5.5
56.9 ±4.7
69.7 ±2.5
73.9 ±5.6
52.1 ±6.7
Table 1: Performance on ALFWorld and WebShop. Locally evaluated RL methods are reported over three decoding seeds. For ALFWorld, we report success rates (%) for six task categories and overall; for WebShop, task score and success rate (%). † indicates results reported in GiGPO ( Feng et al., 2025 ) . The best values are in bold, and the second-best values are underlined.
Figure 5: Sensitivity to the token-residual scale α with Qwen2.5-1.5B-Instruct. Markers and error bars show the mean and sample standard deviation over three decoding seeds.
Figure 6: Turn- and token-level credit in a successful ALFWorld trajectory. Blue, brown, and red denote Akturn , αAtokk,j , and Ak,jVETTA . Only parsed actions are shown; thinking text is abbreviated. When a plotted word spans multiple generated tokens, its values are arithmetic means over those token positions.
Setting
ALFWorld
WebShop
Retained critic blocks
2
2
Tasks × rollouts/update
128×1
128×1
Maximum turns
50
15
PPO minibatch
256
64
Actor / critic LR
10−6/10−5
10−6/10−5
KL loss coefficient
0.01
0.01
Appendix
Table 3: Training settings for VETTA on ALFWorld and WebShop.
Benchmark
Blocks
Value (s)
Update (s)
Local (s)
Share (%)
ALFWorld
2
3.116
16.538
19.696
3.845
ALFWorld
8
10.585
50.460
60.737
9.763
ALFWorld
28
26.700
125.500
151.975
25.634
WebShop
2
1.480
6.546
7.917
4.321
WebShop
8
4.297
18.212
22.455
9.661
WebShop
28
17.579
74.459
92.224
27.783
Appendix
Table 4: Critic timing by backbone depth. All entries are medians over 112 clean training steps. Local time is value inference plus critic update within each step; its share is computed against that step’s total time.
Reinforcement learning with verifiable rewards (RLVR) offers a verifier-bounded performance ceiling for training multi-turn tool-use agents, yet its trajectory-level credit assignment conflates heterogeneous per-turn outcomes into a single reward signal. On-policy distillation provides dense per-token supervision but is either teacher-bounded or prone to gradient concentration collapse. We introduce CrEST, a hierarchical credit assignment framework that retains RL's verifier-bounded ceiling while incorporating dense token-level signals from a privileged self-teacher. CrEST resolves credit at two levels: turn-segmented verified advantages address inter-turn dilution, while entropy-gated self-teacher modulation refines intra-turn token contributions. Experiments on BFCL V3 and WildToolBench show that CrEST consistently outperforms both RL and distillation baselines across two model scales, with the largest gains on long-trajectory and strict session-level metrics. Our work demonstrates that the teacher's role in policy optimization can be reduced from determining update directions to modulating update magnitudes, unlocking dense credit assignment without sacrificing the verifier-bounded ceiling.
Zechuan Wang, Siyuan Lu, Hongxuan Zhang +3
Zhejiang University · AWorld Team, Inclusion AI · Shanghai Innovation Institute +2
Verifier-guided reinforcement learning has become a powerful paradigm for improving LLM reasoning. In multi-turn settings, models receive a verifier score after each turn and iteratively refine their outputs. Although such scores provide dense feedback, they do not directly provide dense credit: a score measures the quality of the current output, while credit should measure how the current turn changes the refinement trajectory. We propose TCPO, a turn-level credit assignment method for verifier-guided multi-turn RL. TCPO casts credit assignment as score-to-credit conversion and constructs turn-level advantages through reference-based comparisons: retrospective credit captures immediate progress and regression relative to the best prior state; hindsight delayed credit identifies non-improving turns with later payoff; and selective fixed-history counterfactual estimation refines high-surprisal turns under the same history. Experiments on math reasoning, code generation, and AppWorld agent tasks show that TCPO improves or matches the strongest baselines across model scales, task domains, and verifier types. TCPO achieves the best or tied-best best-turn Pass@8 on Qwen3-4B and DeepSeek-R1-Distill-Llama-8B, reduces turns to success, and improves multi-turn agent performance. These results highlight score-to-credit conversion as a central ingredient for verifier-guided multi-turn policy optimization.
Multi-turn agents solve complex tasks through extended sequences of tool interactions before producing a final answer, making credit assignment a fundamental challenge during post-training. Outcome rewards provide reliable supervision for short-horizon reasoning, but become sparse and high-variance as trajectories grow to tens or hundreds of tool calls. They can also be misleading: a failed rollout may contain many useful actions that move the agent closer to the goal, yet outcome-only training assigns them the same negative advantage as the eventual mistake. We propose TRACE (Turn-level Reward Assignment via Credit Estimation), a dense credit-assignment method for agentic reinforcement learning. TRACE represents rollouts as state transitions at tool-call boundaries, obtains gold-answer log-probabilities from a frozen reference model, transforms them into log-ratio state values, and derives per-action rewards as Temporal-Difference changes in those values. This requires no additional critic or process-label training, and its one-step log-ratio TD component telescopes across redundant tool calls. On long-horizon complex search, TRACE substantially improves base-model tool-use ability using pure RL, without a cold-start supervised fine-tuning stage, an agentic mid-training stage, or training on live-web data. On the closed-web BrowseComp-Plus benchmark, it raises Qwen3-4B from 7.2 to 35.6 and Qwen3-30B-A3B from 8.4 to 42.6. The learned search behavior also transfers to open-web benchmarks, and the learning curves show earlier improvement and faster convergence during RL training.
Leitian Tao, Baolin Peng, Wenlin Yao +5
University of Wisconsin–Madison · Microsoft Research