LLM agents are emerging as an important paradigm for real-world tasks that require reasoning, tool use, and sequential decision-making. As these agents operate over longer horizons, runtime intervention offers a way to improve reliability without retraining the underlying actor. Effective intervention must provide a useful direction for recovery besides a warning. Existing approaches often rely on an expert solver or a critic that generates task-specific corrections, incurring either the cost of another capable solver or the capacity demands of a task-capable critic. We introduce Comparison-Only Tiny Advisor (COTA) for constructive runtime intervention, which reduces the learned intervention role to local action comparison. A lightweight comparator judges the actor's proposal against available alternatives, and preferred alternatives are returned as non-binding advice for replanning. The comparator is trained from same-prefix counterfactual branches. Across WebShop, ALFWorld, and tau^3-Retail with three LLM actors, COTA instantiated with a 0.5B comparator consistently improves the original actor and achieves the strongest overall performance--cost trade-off among the compared methods. These results suggest that effective runtime intervention need not itself be a task-solving problem: the intervention role can be separated from task solving and handled by a lightweight model specialized for local comparison.
Figures & tables
WebShop
ALFWorld
τ3 -Retail
Actor
Method
Perf. ↑
Cost ↓
Avg. T ↓
Perf. ↑
Cost ↓
Avg. T ↓
Perf. ↑
Cost ↓
Avg. T ↓
Qwen3-8B
Original
39.6
$0.698
1.00 ×
82.8
$0.274
1.00 ×
37.5
$2.00
1.00 ×
Self-Reflection
35.3
$1.54
2.45 ×
83.6
$0.630
1.03 ×
39.2
$4.42
1.66 ×
Asym-AC
10.1
$1.15
2.27 ×
82.8
$0.492
1.03 ×
39.2
$3.20
2.30 ×
AgentRM
40.1
$12.2
16.8 ×
85.1
$3.61
20.6 ×
50.8
$32.8
21.6 ×
COTA (ours)
56.3
$0.883
1.41 ×
90.3
$0.378
1.14 ×
45.0
$2.19
2.04 ×
Table 1: Comparison with baselines. Performance (Perf.) is mean reward on WebShop and success rate on ALFWorld and τ3 -Retail. Best performance is bolded and the second-highest distinct performance is underlined ; the highest inference cost in each actor–environment setting is shown in red . Cost is the estimated API-equivalent inference cost in USD per 100 tasks. Avg. T is the mean task time normalized by the corresponding Original actor.
WebShop
ALFWorld
τ3 -Retail
Objective
Intervention
Qwen3
Qwen3.6
Qwen3
Qwen3.6
Qwen3
Qwen3.6
Absolute Q
Forced
31.5
41.9
2.24
8.96
4.17
5.00
Absolute Q
Constructive
54.9
64.5
57.5
63.4
16.7
17.5
Pairwise comparison
Forced
37.8
60.4
51.5
50.8
16.7
37.5
Pairwise comparison
Constructive
56.3
68.1
90.3
94.0
45.0
65.0
Table 2: Ablation of learning objective and intervention mechanism. WebShop reports mean reward; ALFWorld and τ3 -Retail report success rate. “Constructive” denotes constructive runtime intervention: the proposed action at is withheld only when intervention is triggered, a preferred alternative is returned as advice, and the actor π replans before execution.
Environment
B=0
B=1
B=2
Full
>2 interv.
WebShop
36.6
40.6
44.1
56.3
68.0%
ALFWorld
82.8
83.6
79.9
90.3
37.3%
τ3 -Retail
44.2
43.3
44.2
45.0
19.2%
Table 3: Performance under different episode-level intervention budgets with Qwen3-8B as the actor. B is the maximum number of decision states at which COTA may intervene within an episode; “Full” removes this cap. The “ >2 interv.” column reports the fraction of Full- COTA episodes with more than two interventions.
Appendix figures & tables12 assets
Supplementary material from the paper’s appendix.
Appendix
Environment
Local tasks
DeepSeek tasks
Local repeats
DeepSeek repeats
Max steps
WebShop
500
50
1
1
20
ALFWorld
134
20
1
2
30
τ3 -Retail
40
40
3
3
100
Appendix
Table 4: Evaluation slices. A repeat denotes one complete environment episode with a distinct seeded rollout. DeepSeek uses cost-controlled fixed subsets for WebShop and ALFWorld.
Environment
Physical pairs
Train
Validation
Test
WebShop
35,117
55,856
7,078
7,300
ALFWorld
14,548
22,678
3,222
3,196
τ3 -Retail
5097
8,910
1284
1278
Appendix
Table 5: Comparator data split details.
Configuration
K
R
Mean reward
Success (%)
COTA , cautious
4
4
0.441
18.2
COTA , medium
8
2
0.555
28.4
COTA , active
4
1
0.635
34.2
Appendix
Table 6: Validation performance under different settings of K and R on WebShop.
Environment
Actor condition
Candidate source
WebShop
all actors
environment
ALFWorld
Qwen3-8B
environment
ALFWorld
stronger actors
small LM + environment
τ3 -Retail
all actors
small LM + offline
Appendix
Table 7: Default candidate source by benchmark and actor family.
Environment
Base rollout
Branch rollout
Comparator FT
Total GPU-hours
WebShop
1.48
78.65
1.23
81.36
ALFWorld
1.97
28.39
0.80
31.15
τ3 -Retail
0.27
7.79
0.36
8.42
Appendix
Table 8: Offline cost accounting. The time unit is H200 GPU-hour.
Model
Input ($/M)
Output ($/M)
Qwen3-8B
0.117
0.455
Qwen3.6-35B-A3B
0.050
0.700
DeepSeek-V4-Flash
0.068
0.168
Llama-3.1-8B-Instruct (AgentRM)
0.020
0.040
Appendix
Table 9: API prices used for inference-cost accounting. Prices are in USD per million tokens and correspond to the OpenRouter list prices.
Training target
Pref. consistency
Valid output
End reward
Zero reward
A/B only
39.51%
86.16%
0.5031
37.0%
A/B/T
57.70%
99.53%
0.5442
22.0%
Appendix
Table 10: Comparator-target ablation on 100 held-out WebShop tasks. Both rows use a bidirectional gate and environment-only K=4,R=1 candidates.
Environment
Pairwise ranking acc. (%)
Spearman
BCE
WebShop
69.13
0.618
0.527
ALFWorld
64.23
0.366
0.635
τ3 -Retail
68.20
0.423
0.639
Appendix
Table 11: Validation quality of the three Qwen2.5-0.5B absolute- Q critics.
Qwen3-8B
Qwen3.6
Environment
Forced
Selective
Forced
Selective
WebShop
61.1
46.5
48.8
26.6
ALFWorld
83.6
50.5
87.0
56.5
τ3 -Retail
77.7
38.7
66.0
32.9
Appendix
Table 12: Online intervention frequency (%) for the absolute- Q critic. Forced reports states where argmax replaces the actor proposal; selective reports gate rounds that request actor replanning.
Environment
Actor shift
Same
Tie
Reverse
WebShop
Qwen3-8B → Qwen3.6-35B-A3B
17.0%
79.2%
3.8%
ALFWorld
Qwen3-8B → Qwen3.6-35B-A3B
39.6%
60.4%
0.0%
ALFWorld
Qwen3-8B → DeepSeek-V4-Flash
6.3%
93.8%
0.0%
τ3 -Retail
Qwen3-8B → Qwen3.6-35B-A3B
42.9%
55.7%
1.4%
τ3 -Retail
Qwen3-8B → DeepSeek-V4-Flash
44.3%
52.9%
2.9%
Appendix
Table 13: Changes in pairwise preferences after replacing the Qwen3 continuation actor. Each row is computed over action pairs that are non-tied under Qwen3.
Qwen3-8B
Qwen3.6-35B-A3B
Reference candidate mechanism
Mean reward
Success rate
Mean reward
Success rate
Environment-only
0.563
27.8 %
0.681
56.6 %
Actor sampling + environment
0.602
30.4 %
0.694
57.4 %
Counterfactual prompting + environment
0.579
27.0 %
0.674
56.0 %
Small-LM sampling + environment
0.609
28.8 %
0.696
57.2 %
SFT small-LM sampling + environment
0.583
28.8 %
0.673
54.6 %
Appendix
Table 14: Effect of the reference candidate mechanism on WebShop. Mean reward and exact success rate are reported on the same 500 held-out tasks.
Environment
State representation
WebShop
Task, executed history, current page, executable search/click actions
ALFWorld
Task, executed action–observation history, current observation, admissible actions
τ3 -Retail
Policy and interaction state, including grounded entities, confirmations, tool state, and recent dialogue
Appendix
Table 15: Benchmark-specific state information provided to the comparator.