cs.AIOct 6, 2026

From Uncertainty to Action: Learning to Steer LLM Agents

Authors: Hanwen Li, Jinhao Duan, Guanhua Zhu, Junchi Lu, Bo Shen, Chenxi Yuan, Kaidi Xu

Organizations: New Jersey Institute of Technology · University of North Carolina at Chapel Hill · University of California, Irvine · City University of Hong Kong

Abstract

Steering an LLM agent means deciding whether to correct it, at which step, and with which mechanism. Uncertainty is often used to decide when to correct an agent, but whether it can guide these decisions remains unclear. We steer agent trajectories separately at every non-terminal step with each of four mechanisms and run each continuation to completion. The resulting stepwise outcome table (SOT) holds about 82,000 counterfactual continuations of 1,864 trajectories from three benchmarks and two agents. It shows that uncertainty can identify failing trajectories, but that no single signal reliably locates the step at which steering helps. We therefore propose VoS (Value of Steering), a trajectory-level monitor, offline or online, that learns from SOT the value of steering at each step and decides where to steer by it. A harm-budgeted trigger decides whether to steer, limiting the fraction of successful trajectories that VoS disturbs. VoS improves on unmodified execution in all 12 settings of benchmark, agent, and offline or online use, by 7.8 points on average, and outperforms the strongest of five existing uncertainty-triggered methods in 11, by 2.9 points on average. Ablations show that training on measured outcomes and a tight harm budget are both essential.

Figures & tables

Appendix figures & tables10 assets

Supplementary material from the paper’s appendix.

Appendix

Explore similar work

CardsList
  1. Exploring Agentic Tool-Calling Decisions via Uncertainty-Aligned Reinforcement Learning

    Jun 5, 2026Yijin Zhou, Linqian Zeng, Xiaoya Lu +4Agentic Reinforcement LearningOffline Reinforcement Learning

  2. Calibration Is Not Control: Why LLM-Agent Oversight Needs Intervention

    Jun 19, 2026Chubin Zhang, Zhenglin Wan, Xingrui Yu +5Agentic Control