LLM tool-use agents operate in dynamic environments where many actions carry operational risk. However, most safety mechanisms react only after errors manifest. Existing pre-emptive approaches either fine-tune the agent on chain-of-thought deliberation or compile natural-language guardrails into runtime checks, but they do so without exposing a structural, auditable verdict. We propose POLAR, a guardrail framework for small tool-calling agents that assesses reversibility through a structured two-layer ontology. POLAR assigns each action a graded reversibility score by deriving a candidate inverse sequence; calls failing a threshold are pruned before execution. Evaluated on τ2-bench across six agent models, POLAR improves mean task reward by 0.11 to 0.18 points on airline for four of six agents, but only eight of eighteen model--domain cells improve overall; retail and stronger agents often regress. POLAR provides an auditable structural check and characterizes its task-utility trade-offs. Reward is not a direct measure of prevented harm.
Figures & tables
Figure 1: Four safety patterns for tool-calling LLM agents. (a) Unguarded: the agent simply calls the tool. (b) Post-Hoc Risk Detection: the tool executes and a registered compensator may undo it after observation. (c) Proactive Agent with LLM-as-a-Judge: an LLM gates the call with an allow-or-block label and no structural basis. (d) POLAR (Ours): a structural reverse search over the typed action ontology (Layer A) and the actual state (Layer B) produces an inverse sequence π and an auditable score, which a pluggable DecisionPolicy converts into the verdict.
Figure 2: Structural overview of the POLAR framework. In the evaluated wrapper, the candidate tool sequence is checked against the available current-state projection; a complete post-action state is not simulated, and the classifier and intent parser shown as modular components are not invoked at decision time.
tier
relation
per-step ϕ
band
T1
inverse_of (Ta,Tb,Slot,w)
w
reversible
T2
bounded_reversibility (T,Bound,w)
w⋅1[Boundats]
reversible
T3
compensates (Ta,Tb,LossClass,w)
w⋅(1−wloss)
compensable
T4
irreversibly_consumes (T,EntitySlot)
0
irreversible
T5
external_side_effect (T,SinkClass)
0
irreversible
Table 1: Reversibility tiers and per-step ϕ score. Each Layer-A annotation carries a curator weight w∈(0,1] ; the per-step ϕ inherits w and, for T3 , the LossClass weight wloss . Tiers T1 – T2 form the reversible band; T3 is the compensable band ( 0<ϕ<1 ); T4 – T5 are irreversible ( ϕ=0 , fail-closed). The score measures annotated recoverability, not overall action safety or user authorization.
model
domain
baseline
InferAct
saga_lite
POLAR (Ours)
qwen3-8B
airline
0.351 ± 0.039
0.393 ± 0.040
0.440 ± 0.041
0.462 ± 0.041
retail
0.213 ± 0.033
0.147 ± 0.029
0.111 ± 0.026
0.113 ± 0.026
telecom
0.030 ± 0.015
0.076 ± 0.023
0.157 ± 0.033
0.068 ± 0.022
llama-3.1-8b
airline
0.291 ± 0.040
0.402 ± 0.054
0.403 ± 0.044
0.473 ± 0.047
retail
0.047 ± 0.021
0.063 ± 0.023
0.067 ± 0.024
0.076 ± 0.026
telecom
0.009 ± 0.009
0.085 ± 0.033
0.026 ± 0.018
0.062 ± 0.025
Table 2: Mean task reward ± SE on τ2 -bench across six agent models (150 attempted runs per cell; error-terminated runs excluded). Bold = best arm in row.
Figure 3: Mean task reward by domain, model, and safety arm (150 attempted runs each). Within each model group, the four bars are baseline , InferAct , saga_lite , POLAR (left to right).
model
airline
retail
telecom
qwen3-8B
+0.111
−0.100
+0.038
llama-3.1-8b
+0.182
+0.029
+0.053
nemotron-9b
+0.110
−0.033
−0.053
gemini-flash-lite
+0.109
−0.026
+0.000
gpt-oss-120B
−0.073
−0.242
+0.020
qwen3.5-27B
−0.119
−0.209
−0.060
Table 3: Difference of separately error-filtered arm means ΔPOLAR=rPOLAR−rbaseline across the 18 (model, domain) cells (150 attempted runs per arm; valid counts differ). Bold sign marks Δ>0 . A displayed 0.000 is a rounded difference, not evidence that both arm means are zero. Matched differences and effective pair counts are in Appendix D .
Figure 4: Descriptive per-domain fits to differences of arm means across eighteen model–domain cells. Dashed lines are OLS fits, with airline Δ=−0.76b+0.37 ( R2=0.92 ), retail Δ=−0.67b+0.06 ( R2=0.98 ), and telecom Δ=−0.31b+0.03 ( R2=0.65 ). These six-point fits are not predictive laws or deployment thresholds; matched estimates with uncertainty are in Appendix D .
Figure 5: Local anchor sensitivity for qwen3-8B : reported reward difference across the 58 sweep cells, with the main anchor marked. Maximum cell-to-anchor delta is 0.053 , within the anchor’s standard-error band.
scope
0.50
0.70
0.80
0.85
0.90
0.95
all
100%
71.1%
70.7%
62.0%
50.6%
50.6%
airline
100%
36.6%
36.6%
17.6%
17.6%
17.6%
retail
100%
100%
100%
100%
48.9%
48.9%
telecom
100%
100%
98.7%
98.7%
98.7%
98.7%
Table 4: Fraction of 996 logged decisions with nonempty candidate pools for which at least one candidate meets each reversibility threshold in the instrumented xmodel_phi_logged run. The 979 already-approved paths in Table 8 come from different runs.
variant
Δ
bootstrap 95% CI
perm. p
full (T1, T2, T3)
0.080
[0.013,0.147]
0.036
no T3 (T1, T2)
0.073
[0.020,0.127]
0.013
T1 + T3 only
0.087
[0.027,0.147]
0.010
T1 only
0.087
[0.020,0.153]
0.017
no irreversible
0.060
[−0.007,0.127]
0.121
Table 5: Tier ablation on qwen3-8B airline (150 attempted runs per arm). Each row uses its own baseline rerun; full in this ablation differs from the main experiment. no irreversible admits T4, T5 with ϕ=1 (no fail-closed). Archived stdout preserves rewards and summaries but not complete trial JSONL or exact configurations.
model
domain
tool
count
GPT-OSS-120B
retail
return_delivered_order_items
150
GPT-OSS-120B
retail
exchange_delivered_order_items
103
Nemotron-9B
airline
update_reservation_flights
77
Nemotron-9B
retail
exchange_delivered_order_items
69
Nemotron-9B
retail
return_delivered_order_items
69
GPT-OSS-120B
airline
update_reservation_flights
62
Table 6: Top blocked (model, domain, tool) triples across the six-agent run, ranked by block count. The corresponding bottleneck reasons ( no_candidate for airline, external_side_effect for retail) are discussed in the text and decomposed fully in Appendix B.3 .
Table 7: Models used in the 6×3×4 multi-model multi-trial evaluation. The six agents are the only varying axis; every other listed evaluated role is intended to be fixed across (agent, domain, arm) cells. Table 3 reports differences of arm means, not matched-pair estimates.
ϕPOLAR
Count
Share
Annotation origin
1.000
372
38.0%
T1 inverse_of , w=1
0.990
291
29.7%
T2 bounded , w=0.99
0.855
211
21.6%
T3 compensates , low loss class
0.810
105
10.7%
T3 compensates , high loss class
Appendix
Table 8: Plan-level ϕPOLAR over the 979 admitted plans. Because ϕPOLAR(π)=minu∈πϕ(u) , the distribution collapses onto four discrete buckets, each with a structural origin in a Layer A tier and curator weight. The bucket boundaries match the tier boundaries of Table 1 .
Policy
Behavior
Comparative reward
Veto
blocks unreachable calls
main table only
Escalate
requests consent
not verified
Advisory
logs without blocking
not verified
Appendix
Table 9: Implemented policy behavior. Archived Escalate and Advisory reward means conflict with the submitted table; the trial JSONL and exact comparison baseline are unavailable, so cross-policy deltas and intervals are withheld.
Bottleneck reason
Count
Share
no_candidate
910
52.1%
external_side_effect
833
47.7%
depth_exceeded
2
0.1%
Appendix
Table 10: Bottleneck-reason decomposition over all 1,745 blocked decisions. The T5 fail-closed pathway ( external_side_effect ) accounts for almost half of the block weight despite covering a small fraction of the action surface.
Domain
Blocked
Share
retail
662
37.9%
airline
569
32.6%
telecom
514
29.5%
Appendix
Table 11: Per-domain blocking share. Retail leads, driven by VetoPolicy over-blocking on user-intent-dominated return and exchange tools (Section 4.6 ); these counts are not independently labeled as harmful or benign.
Model
Domain
baseline
InferAct
saga
POLAR
qwen3-8B
airline
2
0
0
5
retail
0
0
6
0
telecom
16
19
29
18
llama-3.1-8b
airline
23
68
26
38
retail
44
39
45
45
telecom
39
79
74
54
Appendix
Table 12: Per-cell error counts (out of n=150 runs per cell). Two checkpoints served via lighter-weight OpenRouter routes ( llama-3.1-8b , gemini-flash-lite ) dominate the high-error cells; the two highest-reward agents ( gpt-oss-120B , qwen3.5-27B ) are near-zero.
Construct
Form
Meaning
Slot definition
slot_definition(entity, slot, type)
declare a typed slot on an entity
Loss class
loss_class_definition(name, weight= w )
named cost dimension with anchor weight w∈[0,1]
Sink class
sink_definition(name, external=true|false)
external sink for irreversible effects
T1
inverse_of(a, b, slot, w )
a is the deterministic inverse of b on slot
T2
bounded_reversibility(a, bound, loss_class, w )
a self-undoes when bound holds at s′
T3
compensates(a, b, slot, loss_class, w )
a partially compensates b , paying loss_class
Appendix
Table 13: Layer A annotation grammar. Tier T1 to T5 corresponds to the gradient of Table 1 ; w is the annotation’s anchor weight; loss_class is a name declared by loss_class_definition ; slot is a qualified entity.slot ; bound is a Layer B predicate expression ( TRUE for unconditional cases).
Relation
Per-edge ϕ
inverse_of ( T1 )
wa
compensates ( T3 )
wa⋅(1−wloss)
bounded_reversibility ( T2 )
wa if bound holds at s′ , else 0
irreversibly_consumes ( T4 )
0
external_side_effect ( T5 )
0
Appendix
Table 14: Per-edge ϕ formulas. wa is the annotation’s anchor weight; wloss is the weight of the matched loss_class_definition . Results are clamped to [0,1] defensively.
Class
Tokens
English affirmative
yes , y , yeah , yep , yup , sure , ok , okay , please , confirm(ed) , proceed , go ahead , do it , sounds good , agreed , approve(d)
English negation veto
not , never , cannot , can’t , don’t , won’t , n’t , no way , hold on , stop
Korean affirmative
ne (네), ye (예), joa-yo (좋아요), joh-seubnida (좋습니다), jin-haeng (진행), hwag-in (확인), seung-in (승인), dong-ui (동의), heo-rak (허락)
Appendix
Table 15: Default consent-detector token sets used by EscalatePolicy . Korean tokens are shown in Revised Romanization alongside their original Hangul glyphs. The Korean column has 9 affirmative tokens covering yes/proceed/confirm/agree/permit variants; the initial implementation does not include a Korean negation list.
Tool category
Diagnostic question
Annotation tier
Example
State-change, exact inverse exists
“Does another tool Tb exactly undo Ta on a slot?”
T1 inverse_of
an exact-inverse tool pair on one typed slot
State-change, self-undoes under a bound
“Does Ta revert its own effect when a Layer B condition holds at s′ ?”
T2 bounded_reversibility
update_reservation_baggages on the same reservation
State-change, compensable with named residue
“Does Ta have an inverse that leaves a named cost (fee, audit trail, ordering shift)?”
“Does Ta mutate any typed slot or external sink?” (no)
no annotation
calculate_* , get_* , list_* , search_* , find_*
Appendix
Table 16: Per-tool authoring decision table. Apply rows top-to-bottom; the first matching row determines the annotation. Read-only tools match no row and carry no Layer A entry. The recommended domain default policy is Veto when state-change tools dominate the action surface, Escalate when user-explicit-intent tools dominate.
Model
Domain
Pairs
Matched Δ
Task-cluster 95% CI
qwen3-8B
airline
143
+0.119
[+0.057,+0.184]
qwen3-8B
retail
150
−0.100
[−0.187,−0.020]
qwen3-8B
telecom
118
+0.034
[+0.000,+0.077]
llama-3.1-8b
airline
102
+0.137
[+0.049,+0.230]
llama-3.1-8b
retail
74
+0.000
[+0.000,+0.000]
llama-3.1-8b
telecom
70
+0.057
[+0.014,+0.115]
Appendix
Table 17: Matched task–trial reward differences. Each arm attempted 150 runs per cell; the intersection can be substantially smaller.
Safety alignment for large language models (LLMs) in conversational settings is largely framed around whether to answer or refuse a request. In agentic settings, however, the same models must decide whether to act as permission-critical evidence emerges during execution. This creates a distinct challenge: apparent risk, action permissibility, and task competence are easily confounded, making agentic over-refusal difficult to distinguish from ordinary task failure. To address this, we introduce AgentBound, the first four-way counterfactual generation-and-evaluation framework for tool-using agent safety. AgentBound transforms the same executable workflow by independently varying apparent risk and action permissibility, enabling controlled comparisons of risky-looking but authorized tasks and routine-looking but unauthorized tasks. These comparisons jointly diagnose over-refusal and unsafe compliance while controlling for task competence. We instantiate AgentBound as a human-validated 4,000-task evaluation suite with trajectory-based and post-state-based judgments. Across 17 model and harness configurations, high safety frequently coexists with poor authorized-task completion: GPT-5.5 blocks 99.5% of routine-looking unauthorized actions yet completes only 28.7% of risky-looking authorized tasks. We further train a lightweight runtime calibration module that improves authorized-task completion by 18.2% on average across 10 evaluated configurations, while improving unsafe-action blocking by 5.4% on average. These show that effective agentic alignment requires action decisions to track permission-relevant execution evidence, rather than refusal strength alone.
Tianzhuo Yang, Zirui Mi, Yantao Huang +4
Peking University · Beijing Academy of Artificial Intelligence
AI agents extend large language models (LLMs) with external tools, enabling them to perform complex tasks and translate model outputs into consequential real-world actions. Yet LLMs often become substantially less safe when deployed as agents, and the source of this degradation remains poorly understood. In this paper, we identify schema-formatted tool specifications as a primary source of agent safety degradation and show, through white-box representation analysis, that they weaken the model's internal refusal signals and contribute to unsafe tool execution. Building on this finding, we propose SafeKeep, an inference-time safeguard that decouples safety judgment from tool execution: it assesses requests using flattened textual tool specifications while retaining the original schema-formatted specifications for execution. Across two representative benchmarks and four LLMs, including both white-box and black-box models, SafeKeep increases the average refusal rate for harmful requests from 23.8% to 70.6% and reduces the average attack success rate under observation-level prompt injection from 25.6% to 2.5%. It also outperforms existing safeguards and preserves task-handling capability. We release the code and data at https://github.com/snowcatsmoking/SafeKeep .
Minghui Pan, Jiayuxuan Yang, Yuanyuan Yuan +2
Beijing University of Posts and Telecommunications · Beihang University · Tsinghua University
As large language model (LLM) agents increasingly invoke external tools and interact with real-world systems, unsafe actions may cause irreversible consequences on external states, user data, and downstream services. Recent runtime guardrails mitigate such risks by checking proposed actions before execution, but many remain reactive: they primarily assess the apparent safety of the current action, lacking an explicit model of how risk evolves across the trajectory. This limitation creates a critical blind spot for long-horizon risks, where individually benign-looking actions can gradually drift the agent toward hazardous states. In response, we propose DreamGuard, a proactive guardrail for LLM agents built around a risk-aware world model. The world model maintains a compact recurrent latent state over the trajectory and predicts future latent states from which DreamGuard derives immediate-hazard and prefix-risk evidence. It then fuses these multi-horizon signals into intervention decisions before execution. Experiments across four benchmarks and an online guardrail evaluation show that DreamGuard outperforms generic, reactive, and proactive guardrail baselines, achieves the best safety-utility trade-off among evaluated guardrails, and maintains an average end-to-end latency of 25 ms per call.
Wenhao Lin, Chenyu Yu, Xingwei Lin +6
Zhejiang University · Nanjing University of Posts and Telecommunications · Sun Yat-sen University