Organizations: New Jersey Institute of Technology · University of North Carolina at Chapel Hill · University of California, Irvine · City University of Hong Kong
Steering an LLM agent means deciding whether to correct it, at which step, and with which mechanism. Uncertainty is often used to decide when to correct an agent, but whether it can guide these decisions remains unclear. We steer agent trajectories separately at every non-terminal step with each of four mechanisms and run each continuation to completion. The resulting stepwise outcome table (SOT) holds about 82,000 counterfactual continuations of 1,864 trajectories from three benchmarks and two agents. It shows that uncertainty can identify failing trajectories, but that no single signal reliably locates the step at which steering helps. We therefore propose VoS (Value of Steering), a trajectory-level monitor, offline or online, that learns from SOT the value of steering at each step and decides where to steer by it. A harm-budgeted trigger decides whether to steer, limiting the fraction of successful trajectories that VoS disturbs. VoS improves on unmodified execution in all 12 settings of benchmark, agent, and offline or online use, by 7.8 points on average, and outperforms the strongest of five existing uncertainty-triggered methods in 11, by 2.9 points on average. Ablations show that training on measured outcomes and a tight harm budget are both essential.
Figures & tables
Figure 1: Overview. (a) An agent with step-level uncertainty. (b) SOT: every step steered with every mechanism and run to the end. (c) VoS: where to steer by a value function learned from SOT, whether by a threshold.
Figure 2: AppWorld, Gemma; Figure 8 shows Qwen. (a) Change in task completion per trajectory when steered at a random step; ticks: mean with 95% bootstrap interval. Steering lowers the score of successful trajectories with every mechanism. (b) Estimators ranked by failure detection and by within-trajectory localization; ρ : Spearman correlation. The best detectors are not the best locators. (c) Per-trajectory Spearman correlation between each estimator and the outcome of memory-based guidance across steps; bar: interquartile range, tick: median. No estimator orders the steps like their outcomes.
References
Steering policies
Unmod.
Always
AUQ
DIAL
Barbi et al.
DS-MCM
ReDAct
VoS (ours)
Offline
AppWorld Gemma
69.0
71.6 ± 1.5
68.1 ± 1.0
68.9 ± 0.3
69.7 ± 1.2
70.4 ± 0.9
69.3 ± 0.7
73.4 ± 1.9
AppWorld Qwen
65.5
55.6 ± 1.6
65.1 ± 0.9
65.4 ± 0.4
65.3 ± 0.7
64.8 ± 0.6
66.0 ± 0.6
66.5 ± 0.8
ALFWorld Gemma
74.5
76.6 ± 2.8
74.7 ± 1.2
80.1 ± 1.8
80.0 ± 1.8
78.4 ± 1.7
76.1 ± 1.0
85.2 ± 1.6
ALFWorld Qwen
52.5
69.0 ± 2.8
62.4 ± 2.5
62.0 ± 3.5
61.0 ± 2.4
57.0 ± 1.6
64.9 ± 2.7
78.4 ± 3.2
Table 1: Main results on each benchmark’s own metric in % (task goal completion on AppWorld, success rate on ALFWorld, score on WebShop), as mean ± standard deviation. Gray: references, not ranked; bold and underline: best and second-best policy in each row.
Figure 3: Final score of VoS against the budget α , on each benchmark’s own metric in %, for Gemma (solid) and Qwen (dashed). Grey lines: unmodified execution; bands: ±1 standard deviation over the splits.
Figure 5
Target
Offline
Online
Failure label
69.3
72.5
DIAL label
70.8
72.1
PRM, labelled
71.6
72.6
PRM, MC rollout
71.4
72.1
Ours
73.4
73.6
Table 2: VoS trained on the targets of prior work: task goal completion in % on AppWorld with Gemma.
Figure 6: Task goal completion against the number of steering events allowed, in real runs (AppWorld, Gemma; line: unmodified execution).
Appendix figures & tables10 assets
Supplementary material from the paper’s appendix.
Appendix
AppWorld
ALFWorld
WebShop
Tasks
258
274
400
Note built from
failed train tasks
150 train games
failed train goals
Action
one Python code block calling app APIs
one text command
search[…] or click[…]
Decoding
temperature 0, seed 100
Appendix
Table 3: Benchmarks and agent settings.
Offline
Online
Model
Gradient-boosted regression trees ( HistGradientBoostingRegressor )
Ridge regression after mean imputation and standardization
Hyperparameters
At most 3 leaves per tree, learning rate 0.05, 200 iterations, at least 20 samples per leaf, ℓ2 regularization 1.0, no early stopping
λ=300
Grid
18 configurations: 3 to 16 leaves, learning rate 0.03 to 0.1, 100 to 400 iterations, 20 or 40 samples per leaf, ℓ2 regularization 0 to 5
7 values, λ∈{10,30,100,…,104}
Appendix
Table 4: Learners and hyperparameters of the monitor, and the grids from which the hyperparameters were chosen.
References
Steering policies (all with memory-based guidance)
Unmod.
Always
AUQ
DIAL
Barbi et al.
DS-MCM
ReDAct
VoS (ours)
Offline
AppWorld Gemma
69.0
71.6 ± 1.5
72.2 ± 1.6
68.9 ± 0.3
69.7 ± 1.2
70.4 ± 0.9
69.5 ± 0.6
73.4 ± 1.9
AppWorld Qwen
65.5
55.6 ± 1.6
64.3 ± 0.7
65.4 ± 0.4
65.3 ± 0.7
64.8 ± 0.6
65.3 ± 0.5
66.5 ± 0.8
ALFWorld Gemma
74.5
76.6 ± 2.8
73.6 ± 0.8
80.1 ± 1.8
80.0 ± 1.8
78.4 ± 1.7
77.4 ± 1.1
85.2 ± 1.6
ALFWorld Qwen
52.5
69.0 ± 2.8
55.3 ± 1.7
62.0 ± 3.5
61.0 ± 2.4
57.0 ± 1.6
63.5 ± 2.6
78.4 ± 3.2
Appendix
Table 5: Table 1 with memory-based guidance for every policy: AUQ and ReDAct steer with it instead of reflection and the stronger model, so that all methods differ only in when and where they steer. Gray: references, not ranked; bold and underline: best and second-best policy in each row.
AppWorld
ALFWorld
WebShop
Offline
Online
Offline
Online
Offline
Online
Memory-based guidance (VoS)
+ .045
+ .046
+ .107
+ .081
+ .055
+ .058
Reflection
+ .008
+ .010
+ .080
+ .018
+ .026
+ .002
Critic
+ .022
+ .002
+ .056
+ .027
+ .051
+ .042
Stronger model
+ .019
+ .015
+ .055
+ .031
+ .046
+ .037
Routed
+ .035
+ .049
+ .092
+ .074
+ .047
+ .055
Appendix
Table 6: Improvement over unmodified execution of the monitor built on each mechanism, and of a routed monitor that chooses at every step the mechanism of highest estimated value (Gemma).
Gemma
Qwen
Tokens
Calls
Steps
Share
Tokens
Calls
Steps
Share
Reflection
2,261
1.00
0.30
4.0%
4,258
1.00
0.48
4.3%
Critic
6,300
3.02
0.84
10.5%
20,094
1.06
2.28
18.8%
Stronger model
6,515
2.04
0.87
9.5%
8,748
2.01
0.99
8.3%
Memory-based guidance
3,589
1.00
0.48
5.1%
6,579
1.00
0.75
6.5%
Appendix
Table 7: Cost of the mechanism’s own call per steering event on AppWorld: prompt and completion tokens, number of calls, the tokens in agent-step equivalents, and their share of the total cost of a steering event.
Reference
Steering policies, net of the rewind
Rewind monitor
AUQ
DIAL
Barbi et al.
DS-MCM
ReDAct
VoS (ours)
Offline
AppWorld Gemma
+ 0.7 ± 0.7
+ 0.8 ± 1.2
+ 1.2 ± 1.7
+ 2.1 ± 1.3
+ 3.4 ± 0.8
+ 1.7 ± 0.8
+ 5.2 ± 1.7
AppWorld Qwen
+ 0.1 ± 1.1
− 0.6 ± 0.5
− 0.4 ± 0.5
− 0.6 ± 0.8
− 1.4 ± 0.9
− 0.1 ± 0.3
+ 1.4 ± 0.7
ALFWorld Gemma
+ 4.6 ± 2.1
− 0.1 ± 0.9
+ 3.6 ± 1.4
+ 3.3 ± 2.0
+ 2.9 ± 1.7
− 1.9 ± 1.3
+ 5.0 ± 1.7
ALFWorld Qwen
+ 9.4 ± 1.9
+ 0.3 ± 0.5
+ 7.4 ± 5.0
+ 4.8 ± 2.7
+ 4.5 ± 2.2
− 1.3 ± 1.1
+ 17.3 ± 4.1
Appendix
Table 8: Change of the benchmark metric net of resampling: for every steered cell, the change produced by a bare rewind at the same step, which continues the trajectory from step t with no message, is subtracted, so a policy is credited only with what its message adds. Metrics as in Table 1 , as mean ± standard deviation. The gray reference is a monitor trained on the rewind alone, which learns where re-running helps. Bold and underline: best and second-best policy in each row; shading from red (loss) to green (gain).
Figure 7: Figure 5 with the cost counted in tokens: those of the steered continuation and the mechanism call, minus, online, those of the replaced remainder.
Gemma
Qwen
Extra pass
re-read
cached
re-read
cached
Probability, entropy
none
0
0
0
0
PMI
action only
150
150
118
118
P(True)
context + action + question
7,559
33
8,857
37
Black-box
context + action + prompt
7,604
75
8,903
75
Agent’s own step
7,529
8,828
Appendix
Table 9: Cost of the uncertainty signals per agent step (AppWorld): tokens read or written in addition to the agent’s own. Re-read: the scoring pass reads the agent’s context again; cached: the context is kept in a prefix cache and only the appended question or prompt is read. Probabilities and entropies come from the pass that generated the action.
Figure 8: Figure 2 for Qwen.
Figure 9: Agreement between mechanisms on failed trajectories (left: Gemma; right: Qwen). Below the diagonal: Spearman correlation of best outcomes across trajectories; above: across steps, averaged over trajectories.
Jun 19, 2026·Chubin Zhang, Zhenglin Wan, Xingrui Yu +5Agentic Control
Nanyang Technological University, Singapore · National University of Singapore, Singapore · CFAR, Agency for Science, Technology and Research, Singapore +2