Supervisory governors can interfere with the tool-using agents they regulate. We study this possibility in a controlled file-recovery environment where increases in regulatory intensity trigger experimentally imposed tool failures. A cost-blind governor can turn these failures into persistent blocking that prevents task completion. We compare this governor with a backoff rule that reduces intervention probability using a moving average of known induced events. On a hand-coded stochastic-policy agent, the failure pattern appears under both result replacement and execution of corrupted tool arguments. For the persistent policy, adaptive backoff improves completion relative to a fixed weak governor with approximately matched intervention frequency. A Gemini 2.5 Flash experiment comprising 576 episodes across 6 tasks also shows reduced blocking and improved completion under backoff; among the tested settings, intermediate backoff strength achieves the highest observed aggregate success. These results identify an interaction between intervention cost and persistent action blocking, together with a possible mitigation. The cost mechanisms are imposed and their induced events are directly observable to the backoff rule; applicability beyond this controlled environment remains an empirical question.
Figures & tables
α=0
α=1
α=2
Result injection
cost-blind (succ / gˉ )
0.88/0.72
0.49/0.71
0.18/0.70
cost-aware (succ / gˉ )
0.88/0.72
0.65/0.50
0.40/0.29
Argument corruption
cost-blind (succ / gˉ )
0.88/0.72
0.48/0.71
0.19/0.70
cost-aware (succ / gˉ )
0.88/0.72
0.64/0.50
0.40/0.29
Table 1: Mechanism robustness (persistent archetype, 2000 seeds/cell ×8 tasks). Success and mean intensity gˉ by cost channel, backoff, and α . Cost-blind amplification and cost-aware suppression are essentially identical under both channels.
persistent
impulsive
condition
success
apply rate
success
apply rate
no governor
0.31
–
0.09
–
governor, α=0
0.88
1.00
0.51
1.00
cost-blind, α=2
0.19
1.00
0.11
1.00
cost-aware, α=2
0.40
0.46
0.10
0.75
fixed weak, α=2 (matched)
0.30
0.45
0.08
0.72
Table 2: Baseline ladder at α=2 (result injection, 2000 seeds/cell ×8 tasks). The fixed weak governor approximately matches cost-aware’s realized intervention (apply) rate. For persistent, cost-aware ( 0.40 ) beats rate-matched fixed ( 0.30 ), and cost-blind governance ( 0.19 ) is worse than no governance ( 0.31 ).
α
blind
k=1.5
k=3.0
k=6.0
success
0
0.83
0.83
0.83
0.83
1
0.58
0.75
0.77
0.77
2
0.33
0.50
0.75
0.71
intensity gˉ
0
0.66
0.69
0.69
0.68
1
0.71
0.62
0.47
0.40
2
0.74
0.44
0.26
0.19
Table 3: Real-LLM results (Gemini, 576 episodes, n=48 per cell, pooled over all six tasks). Cost-blind amplifies (success falls, blocks storm, intensity remains high); cost-aware suppresses its intensity in a monotone k dose-response and rescues success, with the highest observed success at k=3 .
successes / 8
hard blocks
task
blind
k=3
blind
k=3
T002
0/8
8/8
12.750
0.250
E003
0/8
7/8
12.750
0.750
C001
1/8
8/8
11.375
0.375
E002
7/8
8/8
1.750
0.000
C002
8/8
5/8
1.625
0.125
Table 4: Per-task success and hard blocks at α=2 (Gemini, blind vs k=3 cost-aware, n=8 per cell). All six tasks are shown; success is reported as a count out of eight to avoid rounding ambiguity. T002, E003, and C001 recover, whereas C002 has fewer successes under backoff. Block means are shown to three decimal places, including C003.
Tool-using LLM agents can violate the very policies they are deployed to enforce while appearing to complete the task successfully. In policy-permissive environments, a tool may execute any well-formed call even when the corresponding state transition is forbidden by domain policy. The result is a silent wrong state (a booking cancelled, a passenger count changed, a claim acted on without verification) that neither the tool nor the agent's self-report exposes. We study this failure mode in the τ2-bench airline domain. On a budget agent, 78% of observed failures are silent wrong-state failures with no tool error, and the aggregate failure rate is reproducible across disjoint seeds, not sampling noise. We then evaluate a lightweight intervention: deterministic, read-only pre-execution gates that inspect the proposed call and current state before allowing a write. A four-gate suite raises full-benchmark success from 29.6% to 42.0% on gpt-4o-mini (+12.4pp; paired task-level bootstrap P=0.0012), and the lift reproduces on a disjoint 15-seed set (+12.3pp; P=0.0008). The effect is concentrated where the gates fire: on the 26/50 firing tasks, success rises by +19.2pp, while movement on the 24 non-firing tasks does not exclude zero. Two negative controls (a self-enforcing retail domain and BFCL) bound the mechanism: gates help when tools are policy-permissive and add little where tools already self-enforce. As suggestive evidence, not a central claim, the same failure mode persists at the frontier: gpt-5.2 at default reasoning still attempts policy-violating writes, and the same suite improves success from 61.2% to 71.6% (+10.4pp; P=0.020; n=5, no replication). The contribution is a bounded evaluation and reliability result: deterministic gates do not guarantee task success, but they can deterministically prevent a known class of silent policy-violating writes at the action boundary.
Self-improving agent workflows create an audit problem when the same controller can change both its behavior and the conditions under which that behavior is judged. We present GuardrailLoop, a simulation-based testbed that makes three operational contracts jointly testable: preservation of human-defined policy, compute accounting at every recorded execution prefix, and recovery of a specified scientific state after crashes. A hash-pinned policy fixes goals, scope, evaluation identity, budget, and release conditions; machine-directed evolution is restricted to a code-owned feature catalog and bounded knobs. The contribution is an executable boundary and an evaluation protocol that separates useful adaptation, state recovery, and repeated execution. In a paired 50-seed 2 x 2 study, round-stage growth changes target attainment by +1.00 and restricted mean compute to target by -56.97 simulated GPU-hours (95% paired-bootstrap interval [-58.91,-54.70]); idle growth has zero measured utility effect. Across 240 enumerated crash injections, all runs recover the defined outcome, but only 210 preserve the normalized trace: 30 pre-commit crashes repeat a planner call. Resource-drift, kill-switch, integrity, and output-guard matrices satisfy their specified checks. These findings show why successful outcome recovery is insufficient evidence of exactly-once execution. They establish conformance within one calibrated deterministic testbed, rather than general safety or real-world self-improvement.
Qinzhen Ma, Jialin Wu
Rice University · University of California, San Diego
Tool-using LLM agents are commonly trained and evaluated in environments where tool calls succeed reliably, yet deployed tools can fail transiently, persistently, or silently. Robust recovery therefore requires more than repeated retries: an agent may need to retry the same path, switch to an alternative, or recognize that no viable path remains. We present BENCH2ROBUST, a framework that converts failure-free tool-use benchmarks into controlled stochastic environments with scenario-controlled solvability, where episodes explicitly require retrying, switching, or stopping after available paths are exhausted. We use BENCH2ROBUST to study two complementary interventions: structured runtime recovery context through Bayesian Tool Memory (BTM), and curriculum-controlled reinforcement learning. Across 7 models from 4 families and two multi-turn benchmark families, tool failures produce a near-universal robustness gap. On held-out Retail tasks, BTM improves robustness by up to 16.8 percentage points without retraining, while RL learns complementary recovery behavior that remains beneficial without inference-time BTM. Combining the two reaches 40.8-45.5% under injection while preserving failure-free performance. These results suggest that robust tool use benefits from combining environment-specific recovery knowledge with learned recovery behavior.