Organizations: New Jersey Institute of Technology · University of North Carolina at Chapel Hill · University of California, Irvine · City University of Hong Kong
Steering an LLM agent means deciding whether to correct it, at which step, and with which mechanism. Uncertainty is often used to decide when to correct an agent, but whether it can guide these decisions remains unclear. We steer agent trajectories separately at every non-terminal step with each of four mechanisms and run each continuation to completion. The resulting stepwise outcome table (SOT) holds about 82,000 counterfactual continuations of 1,864 trajectories from three benchmarks and two agents. It shows that uncertainty can identify failing trajectories, but that no single signal reliably locates the step at which steering helps. We therefore propose VoS (Value of Steering), a trajectory-level monitor, offline or online, that learns from SOT the value of steering at each step and decides where to steer by it. A harm-budgeted trigger decides whether to steer, limiting the fraction of successful trajectories that VoS disturbs. VoS improves on unmodified execution in all 12 settings of benchmark, agent, and offline or online use, by 7.8 points on average, and outperforms the strongest of five existing uncertainty-triggered methods in 11, by 2.9 points on average. Ablations show that training on measured outcomes and a tight harm budget are both essential.
Figures & tables
Figure 1: Overview. (a) An agent with step-level uncertainty. (b) SOT: every step steered with every mechanism and run to the end. (c) VoS: where to steer by a value function learned from SOT, whether by a threshold.
Figure 2: AppWorld, Gemma; Figure 8 shows Qwen. (a) Change in task completion per trajectory when steered at a random step; ticks: mean with 95% bootstrap interval. Steering lowers the score of successful trajectories with every mechanism. (b) Estimators ranked by failure detection and by within-trajectory localization; ρ : Spearman correlation. The best detectors are not the best locators. (c) Per-trajectory Spearman correlation between each estimator and the outcome of memory-based guidance across steps; bar: interquartile range, tick: median. No estimator orders the steps like their outcomes.
References
Steering policies
Unmod.
Always
AUQ
DIAL
Barbi et al.
DS-MCM
ReDAct
VoS (ours)
Offline
AppWorld Gemma
69.0
71.6 ± 1.5
68.1 ± 1.0
68.9 ± 0.3
69.7 ± 1.2
70.4 ± 0.9
69.3 ± 0.7
73.4 ± 1.9
AppWorld Qwen
65.5
55.6 ± 1.6
65.1 ± 0.9
65.4 ± 0.4
65.3 ± 0.7
64.8 ± 0.6
66.0 ± 0.6
66.5 ± 0.8
ALFWorld Gemma
74.5
76.6 ± 2.8
74.7 ± 1.2
80.1 ± 1.8
80.0 ± 1.8
78.4 ± 1.7
76.1 ± 1.0
85.2 ± 1.6
ALFWorld Qwen
52.5
69.0 ± 2.8
62.4 ± 2.5
62.0 ± 3.5
61.0 ± 2.4
57.0 ± 1.6
64.9 ± 2.7
78.4 ± 3.2
Table 1: Main results on each benchmark’s own metric in % (task goal completion on AppWorld, success rate on ALFWorld, score on WebShop), as mean ± standard deviation. Gray: references, not ranked; bold and underline: best and second-best policy in each row.
Figure 3: Final score of VoS against the budget α , on each benchmark’s own metric in %, for Gemma (solid) and Qwen (dashed). Grey lines: unmodified execution; bands: ±1 standard deviation over the splits.
Figure 5
Target
Offline
Online
Failure label
69.3
72.5
DIAL label
70.8
72.1
PRM, labelled
71.6
72.6
PRM, MC rollout
71.4
72.1
Ours
73.4
73.6
Table 2: VoS trained on the targets of prior work: task goal completion in % on AppWorld with Gemma.
Figure 6: Task goal completion against the number of steering events allowed, in real runs (AppWorld, Gemma; line: unmodified execution).
Appendix figures & tables10 assets
Supplementary material from the paper’s appendix.
Appendix
AppWorld
ALFWorld
WebShop
Tasks
258
274
400
Note built from
failed train tasks
150 train games
failed train goals
Action
one Python code block calling app APIs
one text command
search[…] or click[…]
Decoding
temperature 0, seed 100
Appendix
Table 3: Benchmarks and agent settings.
Offline
Online
Model
Gradient-boosted regression trees ( HistGradientBoostingRegressor )
Ridge regression after mean imputation and standardization
Hyperparameters
At most 3 leaves per tree, learning rate 0.05, 200 iterations, at least 20 samples per leaf, ℓ2 regularization 1.0, no early stopping
λ=300
Grid
18 configurations: 3 to 16 leaves, learning rate 0.03 to 0.1, 100 to 400 iterations, 20 or 40 samples per leaf, ℓ2 regularization 0 to 5
7 values, λ∈{10,30,100,…,104}
Appendix
Table 4: Learners and hyperparameters of the monitor, and the grids from which the hyperparameters were chosen.
References
Steering policies (all with memory-based guidance)
Unmod.
Always
AUQ
DIAL
Barbi et al.
DS-MCM
ReDAct
VoS (ours)
Offline
AppWorld Gemma
69.0
71.6 ± 1.5
72.2 ± 1.6
68.9 ± 0.3
69.7 ± 1.2
70.4 ± 0.9
69.5 ± 0.6
73.4 ± 1.9
AppWorld Qwen
65.5
55.6 ± 1.6
64.3 ± 0.7
65.4 ± 0.4
65.3 ± 0.7
64.8 ± 0.6
65.3 ± 0.5
66.5 ± 0.8
ALFWorld Gemma
74.5
76.6 ± 2.8
73.6 ± 0.8
80.1 ± 1.8
80.0 ± 1.8
78.4 ± 1.7
77.4 ± 1.1
85.2 ± 1.6
ALFWorld Qwen
52.5
69.0 ± 2.8
55.3 ± 1.7
62.0 ± 3.5
61.0 ± 2.4
57.0 ± 1.6
63.5 ± 2.6
78.4 ± 3.2
Appendix
Table 5: Table 1 with memory-based guidance for every policy: AUQ and ReDAct steer with it instead of reflection and the stronger model, so that all methods differ only in when and where they steer. Gray: references, not ranked; bold and underline: best and second-best policy in each row.
AppWorld
ALFWorld
WebShop
Offline
Online
Offline
Online
Offline
Online
Memory-based guidance (VoS)
+ .045
+ .046
+ .107
+ .081
+ .055
+ .058
Reflection
+ .008
+ .010
+ .080
+ .018
+ .026
+ .002
Critic
+ .022
+ .002
+ .056
+ .027
+ .051
+ .042
Stronger model
+ .019
+ .015
+ .055
+ .031
+ .046
+ .037
Routed
+ .035
+ .049
+ .092
+ .074
+ .047
+ .055
Appendix
Table 6: Improvement over unmodified execution of the monitor built on each mechanism, and of a routed monitor that chooses at every step the mechanism of highest estimated value (Gemma).
Gemma
Qwen
Tokens
Calls
Steps
Share
Tokens
Calls
Steps
Share
Reflection
2,261
1.00
0.30
4.0%
4,258
1.00
0.48
4.3%
Critic
6,300
3.02
0.84
10.5%
20,094
1.06
2.28
18.8%
Stronger model
6,515
2.04
0.87
9.5%
8,748
2.01
0.99
8.3%
Memory-based guidance
3,589
1.00
0.48
5.1%
6,579
1.00
0.75
6.5%
Appendix
Table 7: Cost of the mechanism’s own call per steering event on AppWorld: prompt and completion tokens, number of calls, the tokens in agent-step equivalents, and their share of the total cost of a steering event.
Reference
Steering policies, net of the rewind
Rewind monitor
AUQ
DIAL
Barbi et al.
DS-MCM
ReDAct
VoS (ours)
Offline
AppWorld Gemma
+ 0.7 ± 0.7
+ 0.8 ± 1.2
+ 1.2 ± 1.7
+ 2.1 ± 1.3
+ 3.4 ± 0.8
+ 1.7 ± 0.8
+ 5.2 ± 1.7
AppWorld Qwen
+ 0.1 ± 1.1
− 0.6 ± 0.5
− 0.4 ± 0.5
− 0.6 ± 0.8
− 1.4 ± 0.9
− 0.1 ± 0.3
+ 1.4 ± 0.7
ALFWorld Gemma
+ 4.6 ± 2.1
− 0.1 ± 0.9
+ 3.6 ± 1.4
+ 3.3 ± 2.0
+ 2.9 ± 1.7
− 1.9 ± 1.3
+ 5.0 ± 1.7
ALFWorld Qwen
+ 9.4 ± 1.9
+ 0.3 ± 0.5
+ 7.4 ± 5.0
+ 4.8 ± 2.7
+ 4.5 ± 2.2
− 1.3 ± 1.1
+ 17.3 ± 4.1
Appendix
Table 8: Change of the benchmark metric net of resampling: for every steered cell, the change produced by a bare rewind at the same step, which continues the trajectory from step t with no message, is subtracted, so a policy is credited only with what its message adds. Metrics as in Table 1 , as mean ± standard deviation. The gray reference is a monitor trained on the rewind alone, which learns where re-running helps. Bold and underline: best and second-best policy in each row; shading from red (loss) to green (gain).
Figure 7: Figure 5 with the cost counted in tokens: those of the steered continuation and the mechanism call, minus, online, those of the replaced remainder.
Gemma
Qwen
Extra pass
re-read
cached
re-read
cached
Probability, entropy
none
0
0
0
0
PMI
action only
150
150
118
118
P(True)
context + action + question
7,559
33
8,857
37
Black-box
context + action + prompt
7,604
75
8,903
75
Agent’s own step
7,529
8,828
Appendix
Table 9: Cost of the uncertainty signals per agent step (AppWorld): tokens read or written in addition to the agent’s own. Re-read: the scoring pass reads the agent’s context again; cached: the context is kept in a prefix cache and only the appended question or prompt is read. Probabilities and entropies come from the pass that generated the action.
Figure 8: Figure 2 for Qwen.
Figure 9: Agreement between mechanisms on failed trajectories (left: Gemma; right: Qwen). Below the diagonal: Spearman correlation of best outcomes across trajectories; above: across steps, averaged over trajectories.
Large language model (LLM)-based agents often make suboptimal tool-use decisions, including unsupported tool invocation and hallucinated direct responses, which may accumulate errors throughout multi-step interactions. Existing approaches mainly improve these behaviors through inference-time correction or coarse-grained reward signals based on decision outcomes and structured checklists, leaving the uncertainty characteristics of agent decisions underexplored. We observe that decision-oriented reinforcement learning tends to weaken the uncertainty separation between correct and incorrect actions, resulting in overconfident mistakes and weaker exploration signals. Therefore, we propose TRUST, which incorporates uncertainty quantification into reward design as a repulsive force for maintaining uncertainty separation, and labels lightweight key-turn annotations for unified post-training of multi-turn trajectories. Experimental results across diverse tool-use benchmarks show that TRUST consistently enhances both decision quality and agent performance while maintaining more reliable uncertainty estimates during optimization.
Yijin Zhou, Linqian Zeng, Xiaoya Lu +4
Shanghai Jiao Tong University, China · Shanghai Artificial Intelligence Laboratory, China · Shanghai Innovation Institute, China
Uncertainty quantification (UQ) methods for language models are typically evaluated on single-turn outputs, where uncertainty is attached to one generated answer. For LLM agents, however, the unit of observation is an interactive trajectory, where the model can ask clarifying questions, call tools, update state, and make intermediate decisions whose errors propagate to the final outcome. We study whether three common families of single-turn UQ methods transfer to this setting. Across five LLMs and four multi-turn tool-use datasets from BFCL-v4 and τ2-bench, we evaluate white-box scorers based on action-token probabilities, black-box consistency scorers based on resampled trajectories, and reflexive scorers based on model self-assessment of the trajectory. We find that transfer is often useful but uneven. Token-probability scores are highly sensitive to the choice of aggregator used across turns, reflexive scores provide the strongest low-cost baseline in most evaluated settings, and black-box self-consistency is often the strongest UQ family, with trajectory-equivalence and action-set consistency typically ranking highest among its variants. These results suggest that UQ methods developed for single generations should be revalidated at the trajectory level, with careful attention to the consistency measurement, aggregator choice, and computational budget.
Runtime oversight for LLM agents is commonly framed as scalar risk prediction: estimate failure likelihood, confidence, or uncertainty, then intervene once the score crosses a threshold. We argue that this framing targets the wrong object for control. The relevant question is not how likely the agent is to fail if it continues, but whether an available intervention would improve the outcome. Two trajectory prefixes can have the same risk estimate while requiring different actions, because one remains recoverable and the other does not. We formalize this mismatch as target error and identify intervention advantage, the expected utility gain from intervening rather than continuing, as the decision object for oversight. To measure this mismatch, we introduce prefix branching, a same-prefix counterfactual protocol that executes candidate actions from identical trajectory states. Across four benchmarks, action-conditioned control yields regime-dependent gains over scalar routing. In a calibration decomposition, recalibrating the same scalar score improves prediction metrics but leaves control regret unchanged, showing that calibration alone does not repair target error. A simple prefix-only action-conditioned controller substantially reduces regret in the strongest interactive regime, from 0.506 to 0.110 on ALFWorld. Gains shrink when interventions are weak or when scalar routing already preserves intervention-relevant information. These results suggest that LLM-agent oversight should move from calibrated risk scoring toward action-conditioned value estimation.
Chubin Zhang, Zhenglin Wan, Xingrui Yu +5
Nanyang Technological University, Singapore · National University of Singapore, Singapore · CFAR, Agency for Science, Technology and Research, Singapore +2