LLM agents are emerging as an important paradigm for real-world tasks that require reasoning, tool use, and sequential decision-making. As these agents operate over longer horizons, runtime intervention offers a way to improve reliability without retraining the underlying actor. Effective intervention must provide a useful direction for recovery besides a warning. Existing approaches often rely on an expert solver or a critic that generates task-specific corrections, incurring either the cost of another capable solver or the capacity demands of a task-capable critic. We introduce Comparison-Only Tiny Advisor (COTA) for constructive runtime intervention, which reduces the learned intervention role to local action comparison. A lightweight comparator judges the actor's proposal against available alternatives, and preferred alternatives are returned as non-binding advice for replanning. The comparator is trained from same-prefix counterfactual branches. Across WebShop, ALFWorld, and tau^3-Retail with three LLM actors, COTA instantiated with a 0.5B comparator consistently improves the original actor and achieves the strongest overall performance--cost trade-off among the compared methods. These results suggest that effective runtime intervention need not itself be a task-solving problem: the intervention role can be separated from task solving and handled by a lightweight model specialized for local comparison.
Figures & tables
WebShop
ALFWorld
τ3 -Retail
Actor
Method
Perf. ↑
Cost ↓
Avg. T ↓
Perf. ↑
Cost ↓
Avg. T ↓
Perf. ↑
Cost ↓
Avg. T ↓
Qwen3-8B
Original
39.6
$0.698
1.00 ×
82.8
$0.274
1.00 ×
37.5
$2.00
1.00 ×
Self-Reflection
35.3
$1.54
2.45 ×
83.6
$0.630
1.03 ×
39.2
$4.42
1.66 ×
Asym-AC
10.1
$1.15
2.27 ×
82.8
$0.492
1.03 ×
39.2
$3.20
2.30 ×
AgentRM
40.1
$12.2
16.8 ×
85.1
$3.61
20.6 ×
50.8
$32.8
21.6 ×
COTA (ours)
56.3
$0.883
1.41 ×
90.3
$0.378
1.14 ×
45.0
$2.19
2.04 ×
Table 1: Comparison with baselines. Performance (Perf.) is mean reward on WebShop and success rate on ALFWorld and τ3 -Retail. Best performance is bolded and the second-highest distinct performance is underlined ; the highest inference cost in each actor–environment setting is shown in red . Cost is the estimated API-equivalent inference cost in USD per 100 tasks. Avg. T is the mean task time normalized by the corresponding Original actor.
WebShop
ALFWorld
τ3 -Retail
Objective
Intervention
Qwen3
Qwen3.6
Qwen3
Qwen3.6
Qwen3
Qwen3.6
Absolute Q
Forced
31.5
41.9
2.24
8.96
4.17
5.00
Absolute Q
Constructive
54.9
64.5
57.5
63.4
16.7
17.5
Pairwise comparison
Forced
37.8
60.4
51.5
50.8
16.7
37.5
Pairwise comparison
Constructive
56.3
68.1
90.3
94.0
45.0
65.0
Table 2: Ablation of learning objective and intervention mechanism. WebShop reports mean reward; ALFWorld and τ3 -Retail report success rate. “Constructive” denotes constructive runtime intervention: the proposed action at is withheld only when intervention is triggered, a preferred alternative is returned as advice, and the actor π replans before execution.
Environment
B=0
B=1
B=2
Full
>2 interv.
WebShop
36.6
40.6
44.1
56.3
68.0%
ALFWorld
82.8
83.6
79.9
90.3
37.3%
τ3 -Retail
44.2
43.3
44.2
45.0
19.2%
Table 3: Performance under different episode-level intervention budgets with Qwen3-8B as the actor. B is the maximum number of decision states at which COTA may intervene within an episode; “Full” removes this cap. The “ >2 interv.” column reports the fraction of Full- COTA episodes with more than two interventions.
Appendix figures & tables12 assets
Supplementary material from the paper’s appendix.
Appendix
Environment
Local tasks
DeepSeek tasks
Local repeats
DeepSeek repeats
Max steps
WebShop
500
50
1
1
20
ALFWorld
134
20
1
2
30
τ3 -Retail
40
40
3
3
100
Appendix
Table 4: Evaluation slices. A repeat denotes one complete environment episode with a distinct seeded rollout. DeepSeek uses cost-controlled fixed subsets for WebShop and ALFWorld.
Environment
Physical pairs
Train
Validation
Test
WebShop
35,117
55,856
7,078
7,300
ALFWorld
14,548
22,678
3,222
3,196
τ3 -Retail
5097
8,910
1284
1278
Appendix
Table 5: Comparator data split details.
Configuration
K
R
Mean reward
Success (%)
COTA , cautious
4
4
0.441
18.2
COTA , medium
8
2
0.555
28.4
COTA , active
4
1
0.635
34.2
Appendix
Table 6: Validation performance under different settings of K and R on WebShop.
Environment
Actor condition
Candidate source
WebShop
all actors
environment
ALFWorld
Qwen3-8B
environment
ALFWorld
stronger actors
small LM + environment
τ3 -Retail
all actors
small LM + offline
Appendix
Table 7: Default candidate source by benchmark and actor family.
Environment
Base rollout
Branch rollout
Comparator FT
Total GPU-hours
WebShop
1.48
78.65
1.23
81.36
ALFWorld
1.97
28.39
0.80
31.15
τ3 -Retail
0.27
7.79
0.36
8.42
Appendix
Table 8: Offline cost accounting. The time unit is H200 GPU-hour.
Model
Input ($/M)
Output ($/M)
Qwen3-8B
0.117
0.455
Qwen3.6-35B-A3B
0.050
0.700
DeepSeek-V4-Flash
0.068
0.168
Llama-3.1-8B-Instruct (AgentRM)
0.020
0.040
Appendix
Table 9: API prices used for inference-cost accounting. Prices are in USD per million tokens and correspond to the OpenRouter list prices.
Training target
Pref. consistency
Valid output
End reward
Zero reward
A/B only
39.51%
86.16%
0.5031
37.0%
A/B/T
57.70%
99.53%
0.5442
22.0%
Appendix
Table 10: Comparator-target ablation on 100 held-out WebShop tasks. Both rows use a bidirectional gate and environment-only K=4,R=1 candidates.
Environment
Pairwise ranking acc. (%)
Spearman
BCE
WebShop
69.13
0.618
0.527
ALFWorld
64.23
0.366
0.635
τ3 -Retail
68.20
0.423
0.639
Appendix
Table 11: Validation quality of the three Qwen2.5-0.5B absolute- Q critics.
Qwen3-8B
Qwen3.6
Environment
Forced
Selective
Forced
Selective
WebShop
61.1
46.5
48.8
26.6
ALFWorld
83.6
50.5
87.0
56.5
τ3 -Retail
77.7
38.7
66.0
32.9
Appendix
Table 12: Online intervention frequency (%) for the absolute- Q critic. Forced reports states where argmax replaces the actor proposal; selective reports gate rounds that request actor replanning.
Environment
Actor shift
Same
Tie
Reverse
WebShop
Qwen3-8B → Qwen3.6-35B-A3B
17.0%
79.2%
3.8%
ALFWorld
Qwen3-8B → Qwen3.6-35B-A3B
39.6%
60.4%
0.0%
ALFWorld
Qwen3-8B → DeepSeek-V4-Flash
6.3%
93.8%
0.0%
τ3 -Retail
Qwen3-8B → Qwen3.6-35B-A3B
42.9%
55.7%
1.4%
τ3 -Retail
Qwen3-8B → DeepSeek-V4-Flash
44.3%
52.9%
2.9%
Appendix
Table 13: Changes in pairwise preferences after replacing the Qwen3 continuation actor. Each row is computed over action pairs that are non-tied under Qwen3.
Qwen3-8B
Qwen3.6-35B-A3B
Reference candidate mechanism
Mean reward
Success rate
Mean reward
Success rate
Environment-only
0.563
27.8 %
0.681
56.6 %
Actor sampling + environment
0.602
30.4 %
0.694
57.4 %
Counterfactual prompting + environment
0.579
27.0 %
0.674
56.0 %
Small-LM sampling + environment
0.609
28.8 %
0.696
57.2 %
SFT small-LM sampling + environment
0.583
28.8 %
0.673
54.6 %
Appendix
Table 14: Effect of the reference candidate mechanism on WebShop. Mean reward and exact success rate are reported on the same 500 held-out tasks.
Environment
State representation
WebShop
Task, executed history, current page, executable search/click actions
ALFWorld
Task, executed action–observation history, current observation, admissible actions
τ3 -Retail
Policy and interaction state, including grounded entities, confirmations, tool state, and recent dialogue
Appendix
Table 15: Benchmark-specific state information provided to the comparator.
Large language model agents rely on external harnesses to pass information between the model and its environment and to recover from execution errors. Yet recovery is usually judged only by average task success. This hides an important tension. The same operation can rescue a failing trajectory or disrupt one that would otherwise succeed. We frame recovery as a causal decision problem. Starting from the same execution state, we compare what happens with and without recovery, separate rescue from harm, and study how the value of recovery changes over time. We then introduce the Causal Intervention Router (CIR), a lightweight policy that uses information available before recovery to decide when intervention is worthwhile. On long-horizon ALFWorld tasks with Qwen3-14B, CIR raises success from 70.33% to 73.33%, a gain of 3.00 percentage points. It leaves all evaluated trajectories with correct observations untouched. Additional controls show that the benefit of recovery cannot be explained solely by the new observation returned by the environment. These results provide a practical way to evaluate recovery and apply it selectively.
Shuyao Xiao, Shengling Wang, Xuan Chen +7
School of Artificial Intelligence, Beijing Normal University · Ke Holdings
Large language model agents interleave reasoning, action selection, and observation to solve sequential decision-making tasks. In deployed settings where agents repeatedly handle related multi-step tasks, small action-selection errors can accumulate into wasted tool calls, latency, and reduced reliability. Despite this need for deployment-time improvement, existing inference-time adaptation methods for LLM agents mainly rely on prompting or retrieval, which influence behavior indirectly through context manipulation. For ReAct-style agents, such approaches do not expose an explicit decision layer that can score candidate actions, represent uncertainty, or be updated online from action-level feedback. As a result, they provide limited support for trackable, fine-grained, and uncertainty-aware adaptation during deployment. We propose OLIVIA, an inference-time action adaptation framework for ReAct-style agents. OLIVIA models the LLM's final action-selection layer as a contextual linear bandit over candidate actions, with frozen hidden states as decision contexts. This choice is particularly suitable for deployment because it adapts behavior directly at the action-selection interface, preserves the underlying reasoning process, and provides explicit uncertainty estimates and lightweight online updates from action-level feedback. With upper-confidence-bound exploration, OLIVIA improves the policy sample-efficiently with minimal computational overhead. We instantiate OLIVIA on four benchmarks and show that it consistently improves task performance over static ReAct and prompt-based inference-time baselines. Our results suggest that explicit online decision layers provide an effective alternative to purely prompt- or retrieval-based adaptation for LLM agents during deployment.
Sheldon Yu, Junda Wu, Xintong Li +6
UC San Diego · University of Illinois at Urbana-Champaign · Adobe Research
Large language models (LLMs) are increasingly used as tutors and thought partners, helping users reason through problems. While guidance from AI assistants can scaffold thinking and foster learning, such benefits depend on how they help--for instance, intervening too early or too frequently may hinder true learning and cognitive engagement. Yet how AI systems navigate intervention decisions during problem-solving remains poorly understood. Here, we introduce Int-Bench, a simulation-based benchmark for evaluating LLM interventions during learning. Int-Bench simulates a "student" solving a problem while a "teacher" monitors the student's reasoning and decides whether, when, and how to intervene. Across three domains--code debugging, mathematics, and brain teasers--we evaluate LLM teachers on the frequency and timing of interventions, as well as their impact on both immediate task success and generalization to new problems. We also compare LLMs to humans, finding that LLMs intervene more frequently and earlier than humans. Moreover, in contrast to humans, they tend to provide complete solutions rather than targeted hints. These findings suggest that current LLM assistants often optimize for short-term success rather than supporting the reasoning processes needed for deeper learning and long-term success.
Verona Teo, Raghav Jain, Tobias Gerstenberg +1
1Stanford University · University of California, San Diego · University of Washington