Tool-using agents are continually updated with new interaction data. After each policy update, however, previously estimated action credits may become stale. Recomputing them from scratch can require many additional tool calls and environment interactions, making repeated updates increasingly expensive. We ask a simple question: when does historical action credit actually need to be updated? Our key observation is that a change in action value does not necessarily imply a change in the decision. Historical credit can still be useful as long as policy-induced drift is too small to overturn the existing action ranking. Building on this idea, we introduce pairwise branch sensitivity to capture how strongly a policy update affects the downstream regions that distinguish two candidate actions. We then derive a first-order anchored credit-transport estimator that updates historical credit using old interventional trajectories, and propose a Decision-Sufficient Credit Gate (DSC-Gate) that chooses whether to reuse, transport, or resample credit. Experiments show that branch sensitivity explains credit drift substantially better than global policy distance. With sufficient historical data, credit transport reduces estimation error, while its benefit to decision making is concentrated on updates that affect action-distinguishing branches. On a fully independent test set, DSC-Gate changes mean regret by only +0.00004 relative to a gap-based gate while reducing mean new tool steps from 472 to 286, a 39.4% reduction. We observe the same pattern after a real tool-agent parameter update. Overall, our results show that agents do not need to recompute action credit after every policy update: much of the historical evidence can be reused or cheaply corrected, reducing the additional interaction required to keep action decisions up to date.
Figures & tables
Figure 1: Deciding when historical action credit must be refreshed. Left to right: after a policy update μ→π , old interventional trajectories Dμ collected under the previous policy are reused to compare two candidate root actions a and b . Pairwise branch sensitivity G(a,b) weights the update by the visitation difference dμ,a−dμ,b , so changes in states that both branches visit cancel in the action difference, and only changes in branch-differential states drive relative drift. First-order anchored transport then corrects the historical gap as cT=cR+δ from the same trajectories, without new execution. DSC-Gate combines cR , δ , and empirically calibrated radii rR,rΔ,rT to resolve each leader-versus-competitor comparison by reuse, transport, or refresh; only refresh executes the target policy in the environment. The bars show how the gate shifts toward refresh when the update is concentrated on branch-differential states.
Figure 2: Credit drift and ranking stability. Each panel contains all 864 action pairs from the 288 L2–L3 confirmation tasks for a different update condition. Signs are oriented by old credit. Highlighted points are true ranking reversals; the annotated rate additionally requires the absolute new credit difference to exceed the prespecified relevance threshold.
Figure 3: Branch sensitivity and credit drift in L1. Panel A shows the gain in Spearman correlation from replacing global KL with pairwise KL. Panel B shows the additional gain from directed branch sensitivity over pairwise KL. Each hexagon aggregates one or more confirmation tasks; the dashed line marks zero gain.
Baseline
Baseline AUC
Difference
Simultaneous interval
Adaptive DR
0.008462
-0.001583
[-0.002651, -0.000544]
Occupancy-sensitive
0.006525
+0.000354
[-0.000147, +0.000891]
Gap-based
0.006494
+0.000385
[-0.000124, +0.000950]
Table 1: Primary L3 decision comparisons. Differences are transport-adaptive minus the baseline; lower is better. Intervals are simultaneous for the three prespecified comparisons.
Figure 4: From credit correction to decision benefit. Panel A reports the relative reduction in pairwise credit MSE from transport over direct reuse across update conditions and old-data budgets. Panel B reports the corresponding decision benefit across old-data and new-execution budgets. Numerical improvement is most useful when the update can change the action ranking.
Method
Mean regret
P(R>0.02)
New steps
Any refresh
DSC-Gate
0.00673
2.9%
286
38.9%
Gap-Gate
0.00669
3.0%
472
63.5%
WIS-Gate
0.00734
3.8%
521
69.4%
DR-Gate
0.00812
4.6%
558
72.6%
Reuse only
0.01284
8.1%
0
0%
Transport only
0.00803
4.7%
0
0%
Table 2: Gate results on the independent L4 test set.