Tool-using agents are continually updated with new interaction data. After each policy update, however, previously estimated action credits may become stale. Recomputing them from scratch can require many additional tool calls and environment interactions, making repeated updates increasingly expensive. We ask a simple question: when does historical action credit actually need to be updated? Our key observation is that a change in action value does not necessarily imply a change in the decision. Historical credit can still be useful as long as policy-induced drift is too small to overturn the existing action ranking. Building on this idea, we introduce pairwise branch sensitivity to capture how strongly a policy update affects the downstream regions that distinguish two candidate actions. We then derive a first-order anchored credit-transport estimator that updates historical credit using old interventional trajectories, and propose a Decision-Sufficient Credit Gate (DSC-Gate) that chooses whether to reuse, transport, or resample credit. Experiments show that branch sensitivity explains credit drift substantially better than global policy distance. With sufficient historical data, credit transport reduces estimation error, while its benefit to decision making is concentrated on updates that affect action-distinguishing branches. On a fully independent test set, DSC-Gate changes mean regret by only +0.00004 relative to a gap-based gate while reducing mean new tool steps from 472 to 286, a 39.4% reduction. We observe the same pattern after a real tool-agent parameter update. Overall, our results show that agents do not need to recompute action credit after every policy update: much of the historical evidence can be reused or cheaply corrected, reducing the additional interaction required to keep action decisions up to date.
Figures & tables
Figure 1: Deciding when historical action credit must be refreshed. Left to right: after a policy update μ→π , old interventional trajectories Dμ collected under the previous policy are reused to compare two candidate root actions a and b . Pairwise branch sensitivity G(a,b) weights the update by the visitation difference dμ,a−dμ,b , so changes in states that both branches visit cancel in the action difference, and only changes in branch-differential states drive relative drift. First-order anchored transport then corrects the historical gap as cT=cR+δ from the same trajectories, without new execution. DSC-Gate combines cR , δ , and empirically calibrated radii rR,rΔ,rT to resolve each leader-versus-competitor comparison by reuse, transport, or refresh; only refresh executes the target policy in the environment. The bars show how the gate shifts toward refresh when the update is concentrated on branch-differential states.
Figure 2: Credit drift and ranking stability. Each panel contains all 864 action pairs from the 288 L2–L3 confirmation tasks for a different update condition. Signs are oriented by old credit. Highlighted points are true ranking reversals; the annotated rate additionally requires the absolute new credit difference to exceed the prespecified relevance threshold.
Figure 3: Branch sensitivity and credit drift in L1. Panel A shows the gain in Spearman correlation from replacing global KL with pairwise KL. Panel B shows the additional gain from directed branch sensitivity over pairwise KL. Each hexagon aggregates one or more confirmation tasks; the dashed line marks zero gain.
Baseline
Baseline AUC
Difference
Simultaneous interval
Adaptive DR
0.008462
-0.001583
[-0.002651, -0.000544]
Occupancy-sensitive
0.006525
+0.000354
[-0.000147, +0.000891]
Gap-based
0.006494
+0.000385
[-0.000124, +0.000950]
Table 1: Primary L3 decision comparisons. Differences are transport-adaptive minus the baseline; lower is better. Intervals are simultaneous for the three prespecified comparisons.
Figure 4: From credit correction to decision benefit. Panel A reports the relative reduction in pairwise credit MSE from transport over direct reuse across update conditions and old-data budgets. Panel B reports the corresponding decision benefit across old-data and new-execution budgets. Numerical improvement is most useful when the update can change the action ranking.
Method
Mean regret
P(R>0.02)
New steps
Any refresh
DSC-Gate
0.00673
2.9%
286
38.9%
Gap-Gate
0.00669
3.0%
472
63.5%
WIS-Gate
0.00734
3.8%
521
69.4%
DR-Gate
0.00812
4.6%
558
72.6%
Reuse only
0.01284
8.1%
0
0%
Transport only
0.00803
4.7%
0
0%
Table 2: Gate results on the independent L4 test set.
Credit assignment in multi-turn agent reinforcement learning operates at two levels: assigning trajectory-level credit to actions and distributing each action's credit across its tokens. In this paper, we introduce FACTOR, which separates these decisions. FACTOR uses checkpoint-calibrated TD residuals to assign per-action credits that telescope to the trajectory advantage, and feedback-conditioned teacher-student likelihood gaps to allocate each credit across the realized action tokens. Per-action normalization preserves the action-average coefficient and prevents token-level sign flips. We pair this construction with an action-mean reduction, removing the implicit dependence of an action's scalar surrogate weight on its token length. At the behavior policy and before clipping, each action's inner action-mean surrogate equals its TD credit. FACTOR consistently improves over competitive baselines across ALFWorld, WebShop, and ScienceWorld, with every environment-seed comparison favoring FACTOR and the largest gains emerging on the longest-horizon environment. The same hyperparameters transfer without retuning to a larger backbone and to a different model family. Ablations identify TD action credit as the dominant driver of the improvement, with hindsight token allocation contributing complementary gains.
Lichao Ma, Yang Sun, Shuaitao Zhao +9
1Peking University · 2Meituan LongCat Interaction Team · 3Fudan University +4
Multi-turn agents solve complex tasks through extended sequences of tool interactions before producing a final answer, making credit assignment a fundamental challenge during post-training. Outcome rewards provide reliable supervision for short-horizon reasoning, but become sparse and high-variance as trajectories grow to tens or hundreds of tool calls. They can also be misleading: a failed rollout may contain many useful actions that move the agent closer to the goal, yet outcome-only training assigns them the same negative advantage as the eventual mistake. We propose TRACE (Turn-level Reward Assignment via Credit Estimation), a dense credit-assignment method for agentic reinforcement learning. TRACE represents rollouts as state transitions at tool-call boundaries, obtains gold-answer log-probabilities from a frozen reference model, transforms them into log-ratio state values, and derives per-action rewards as Temporal-Difference changes in those values. This requires no additional critic or process-label training, and its one-step log-ratio TD component telescopes across redundant tool calls. On long-horizon complex search, TRACE substantially improves base-model tool-use ability using pure RL, without a cold-start supervised fine-tuning stage, an agentic mid-training stage, or training on live-web data. On the closed-web BrowseComp-Plus benchmark, it raises Qwen3-4B from 7.2 to 35.6 and Qwen3-30B-A3B from 8.4 to 42.6. The learned search behavior also transfers to open-web benchmarks, and the learning curves show earlier improvement and faster convergence during RL training.
Leitian Tao, Baolin Peng, Wenlin Yao +5
University of Wisconsin–Madison · Microsoft Research
Long-horizon tool-use reinforcement learning learns from outcome verification, but trajectory-level advantages are broadcast over reasoning, API, and answer tokens. Direct self-distillation can supply a denser signal, but in our experiments it can also destroy tool use by rehearsing teacher behavior without identifying which actions the verifier rewards. We introduce Sibling-Guided Credit Distillation (SGCD), which uses distillation for bounded credit weighting rather than as a competing actor loss. Dynamic sampling produces mixed successful and failed sibling rollouts; an external LLM summarizes their contrast into a training-only credit reference; and detached teacher/student divergence reshapes GRPO token advantages. The deployed student receives only the clean task prompt. Across AppWorld and tau^3-airline, SGCD reports higher held-out point estimates than GRPO-family comparators: AppWorld TGC improves from 42.9 to 45.6 on test_normal and from 24.7 to 27.0 on test_challenge, and tau^3-airline held-out evaluator score improves from 0.583 to 0.602. These results support a narrow design rule for long-horizon tool-use agents: use distillation to guide credit assignment while keeping policy gradient in charge of the actor update.
Tianyu Ding, Jianhong Xin, Juan Pablo De la Cruz Weinstein