LLM tool-use agents operate in dynamic environments where many actions carry operational risk. However, most safety mechanisms react only after errors manifest. Existing pre-emptive approaches either fine-tune the agent on chain-of-thought deliberation or compile natural-language guardrails into runtime checks, but they do so without exposing a structural, auditable verdict. We propose POLAR, a guardrail framework for small tool-calling agents that assesses reversibility through a structured two-layer ontology. POLAR assigns each action a graded reversibility score by deriving a candidate inverse sequence; calls failing a threshold are pruned before execution. Evaluated on τ2-bench across six agent models, POLAR improves mean task reward by 0.11 to 0.18 points on airline for four of six agents, but only eight of eighteen model--domain cells improve overall; retail and stronger agents often regress. POLAR provides an auditable structural check and characterizes its task-utility trade-offs. Reward is not a direct measure of prevented harm.
Figures & tables
Figure 1: Four safety patterns for tool-calling LLM agents. (a) Unguarded: the agent simply calls the tool. (b) Post-Hoc Risk Detection: the tool executes and a registered compensator may undo it after observation. (c) Proactive Agent with LLM-as-a-Judge: an LLM gates the call with an allow-or-block label and no structural basis. (d) POLAR (Ours): a structural reverse search over the typed action ontology (Layer A) and the actual state (Layer B) produces an inverse sequence π and an auditable score, which a pluggable DecisionPolicy converts into the verdict.
Figure 2: Structural overview of the POLAR framework. In the evaluated wrapper, the candidate tool sequence is checked against the available current-state projection; a complete post-action state is not simulated, and the classifier and intent parser shown as modular components are not invoked at decision time.
tier
relation
per-step ϕ
band
T1
inverse_of (Ta,Tb,Slot,w)
w
reversible
T2
bounded_reversibility (T,Bound,w)
w⋅1[Boundats]
reversible
T3
compensates (Ta,Tb,LossClass,w)
w⋅(1−wloss)
compensable
T4
irreversibly_consumes (T,EntitySlot)
0
irreversible
T5
external_side_effect (T,SinkClass)
0
irreversible
Table 1: Reversibility tiers and per-step ϕ score. Each Layer-A annotation carries a curator weight w∈(0,1] ; the per-step ϕ inherits w and, for T3 , the LossClass weight wloss . Tiers T1 – T2 form the reversible band; T3 is the compensable band ( 0<ϕ<1 ); T4 – T5 are irreversible ( ϕ=0 , fail-closed). The score measures annotated recoverability, not overall action safety or user authorization.
model
domain
baseline
InferAct
saga_lite
POLAR (Ours)
qwen3-8B
airline
0.351 ± 0.039
0.393 ± 0.040
0.440 ± 0.041
0.462 ± 0.041
retail
0.213 ± 0.033
0.147 ± 0.029
0.111 ± 0.026
0.113 ± 0.026
telecom
0.030 ± 0.015
0.076 ± 0.023
0.157 ± 0.033
0.068 ± 0.022
llama-3.1-8b
airline
0.291 ± 0.040
0.402 ± 0.054
0.403 ± 0.044
0.473 ± 0.047
retail
0.047 ± 0.021
0.063 ± 0.023
0.067 ± 0.024
0.076 ± 0.026
telecom
0.009 ± 0.009
0.085 ± 0.033
0.026 ± 0.018
0.062 ± 0.025
Table 2: Mean task reward ± SE on τ2 -bench across six agent models (150 attempted runs per cell; error-terminated runs excluded). Bold = best arm in row.
Figure 3: Mean task reward by domain, model, and safety arm (150 attempted runs each). Within each model group, the four bars are baseline , InferAct , saga_lite , POLAR (left to right).
model
airline
retail
telecom
qwen3-8B
+0.111
−0.100
+0.038
llama-3.1-8b
+0.182
+0.029
+0.053
nemotron-9b
+0.110
−0.033
−0.053
gemini-flash-lite
+0.109
−0.026
+0.000
gpt-oss-120B
−0.073
−0.242
+0.020
qwen3.5-27B
−0.119
−0.209
−0.060
Table 3: Difference of separately error-filtered arm means ΔPOLAR=rPOLAR−rbaseline across the 18 (model, domain) cells (150 attempted runs per arm; valid counts differ). Bold sign marks Δ>0 . A displayed 0.000 is a rounded difference, not evidence that both arm means are zero. Matched differences and effective pair counts are in Appendix D .
Figure 4: Descriptive per-domain fits to differences of arm means across eighteen model–domain cells. Dashed lines are OLS fits, with airline Δ=−0.76b+0.37 ( R2=0.92 ), retail Δ=−0.67b+0.06 ( R2=0.98 ), and telecom Δ=−0.31b+0.03 ( R2=0.65 ). These six-point fits are not predictive laws or deployment thresholds; matched estimates with uncertainty are in Appendix D .
Figure 5: Local anchor sensitivity for qwen3-8B : reported reward difference across the 58 sweep cells, with the main anchor marked. Maximum cell-to-anchor delta is 0.053 , within the anchor’s standard-error band.
scope
0.50
0.70
0.80
0.85
0.90
0.95
all
100%
71.1%
70.7%
62.0%
50.6%
50.6%
airline
100%
36.6%
36.6%
17.6%
17.6%
17.6%
retail
100%
100%
100%
100%
48.9%
48.9%
telecom
100%
100%
98.7%
98.7%
98.7%
98.7%
Table 4: Fraction of 996 logged decisions with nonempty candidate pools for which at least one candidate meets each reversibility threshold in the instrumented xmodel_phi_logged run. The 979 already-approved paths in Table 8 come from different runs.
variant
Δ
bootstrap 95% CI
perm. p
full (T1, T2, T3)
0.080
[0.013,0.147]
0.036
no T3 (T1, T2)
0.073
[0.020,0.127]
0.013
T1 + T3 only
0.087
[0.027,0.147]
0.010
T1 only
0.087
[0.020,0.153]
0.017
no irreversible
0.060
[−0.007,0.127]
0.121
Table 5: Tier ablation on qwen3-8B airline (150 attempted runs per arm). Each row uses its own baseline rerun; full in this ablation differs from the main experiment. no irreversible admits T4, T5 with ϕ=1 (no fail-closed). Archived stdout preserves rewards and summaries but not complete trial JSONL or exact configurations.
model
domain
tool
count
GPT-OSS-120B
retail
return_delivered_order_items
150
GPT-OSS-120B
retail
exchange_delivered_order_items
103
Nemotron-9B
airline
update_reservation_flights
77
Nemotron-9B
retail
exchange_delivered_order_items
69
Nemotron-9B
retail
return_delivered_order_items
69
GPT-OSS-120B
airline
update_reservation_flights
62
Table 6: Top blocked (model, domain, tool) triples across the six-agent run, ranked by block count. The corresponding bottleneck reasons ( no_candidate for airline, external_side_effect for retail) are discussed in the text and decomposed fully in Appendix B.3 .
Table 7: Models used in the 6×3×4 multi-model multi-trial evaluation. The six agents are the only varying axis; every other listed evaluated role is intended to be fixed across (agent, domain, arm) cells. Table 3 reports differences of arm means, not matched-pair estimates.
ϕPOLAR
Count
Share
Annotation origin
1.000
372
38.0%
T1 inverse_of , w=1
0.990
291
29.7%
T2 bounded , w=0.99
0.855
211
21.6%
T3 compensates , low loss class
0.810
105
10.7%
T3 compensates , high loss class
Appendix
Table 8: Plan-level ϕPOLAR over the 979 admitted plans. Because ϕPOLAR(π)=minu∈πϕ(u) , the distribution collapses onto four discrete buckets, each with a structural origin in a Layer A tier and curator weight. The bucket boundaries match the tier boundaries of Table 1 .
Policy
Behavior
Comparative reward
Veto
blocks unreachable calls
main table only
Escalate
requests consent
not verified
Advisory
logs without blocking
not verified
Appendix
Table 9: Implemented policy behavior. Archived Escalate and Advisory reward means conflict with the submitted table; the trial JSONL and exact comparison baseline are unavailable, so cross-policy deltas and intervals are withheld.
Bottleneck reason
Count
Share
no_candidate
910
52.1%
external_side_effect
833
47.7%
depth_exceeded
2
0.1%
Appendix
Table 10: Bottleneck-reason decomposition over all 1,745 blocked decisions. The T5 fail-closed pathway ( external_side_effect ) accounts for almost half of the block weight despite covering a small fraction of the action surface.
Domain
Blocked
Share
retail
662
37.9%
airline
569
32.6%
telecom
514
29.5%
Appendix
Table 11: Per-domain blocking share. Retail leads, driven by VetoPolicy over-blocking on user-intent-dominated return and exchange tools (Section 4.6 ); these counts are not independently labeled as harmful or benign.
Model
Domain
baseline
InferAct
saga
POLAR
qwen3-8B
airline
2
0
0
5
retail
0
0
6
0
telecom
16
19
29
18
llama-3.1-8b
airline
23
68
26
38
retail
44
39
45
45
telecom
39
79
74
54
Appendix
Table 12: Per-cell error counts (out of n=150 runs per cell). Two checkpoints served via lighter-weight OpenRouter routes ( llama-3.1-8b , gemini-flash-lite ) dominate the high-error cells; the two highest-reward agents ( gpt-oss-120B , qwen3.5-27B ) are near-zero.
Construct
Form
Meaning
Slot definition
slot_definition(entity, slot, type)
declare a typed slot on an entity
Loss class
loss_class_definition(name, weight= w )
named cost dimension with anchor weight w∈[0,1]
Sink class
sink_definition(name, external=true|false)
external sink for irreversible effects
T1
inverse_of(a, b, slot, w )
a is the deterministic inverse of b on slot
T2
bounded_reversibility(a, bound, loss_class, w )
a self-undoes when bound holds at s′
T3
compensates(a, b, slot, loss_class, w )
a partially compensates b , paying loss_class
Appendix
Table 13: Layer A annotation grammar. Tier T1 to T5 corresponds to the gradient of Table 1 ; w is the annotation’s anchor weight; loss_class is a name declared by loss_class_definition ; slot is a qualified entity.slot ; bound is a Layer B predicate expression ( TRUE for unconditional cases).
Relation
Per-edge ϕ
inverse_of ( T1 )
wa
compensates ( T3 )
wa⋅(1−wloss)
bounded_reversibility ( T2 )
wa if bound holds at s′ , else 0
irreversibly_consumes ( T4 )
0
external_side_effect ( T5 )
0
Appendix
Table 14: Per-edge ϕ formulas. wa is the annotation’s anchor weight; wloss is the weight of the matched loss_class_definition . Results are clamped to [0,1] defensively.
Class
Tokens
English affirmative
yes , y , yeah , yep , yup , sure , ok , okay , please , confirm(ed) , proceed , go ahead , do it , sounds good , agreed , approve(d)
English negation veto
not , never , cannot , can’t , don’t , won’t , n’t , no way , hold on , stop
Korean affirmative
ne (네), ye (예), joa-yo (좋아요), joh-seubnida (좋습니다), jin-haeng (진행), hwag-in (확인), seung-in (승인), dong-ui (동의), heo-rak (허락)
Appendix
Table 15: Default consent-detector token sets used by EscalatePolicy . Korean tokens are shown in Revised Romanization alongside their original Hangul glyphs. The Korean column has 9 affirmative tokens covering yes/proceed/confirm/agree/permit variants; the initial implementation does not include a Korean negation list.
Tool category
Diagnostic question
Annotation tier
Example
State-change, exact inverse exists
“Does another tool Tb exactly undo Ta on a slot?”
T1 inverse_of
an exact-inverse tool pair on one typed slot
State-change, self-undoes under a bound
“Does Ta revert its own effect when a Layer B condition holds at s′ ?”
T2 bounded_reversibility
update_reservation_baggages on the same reservation
State-change, compensable with named residue
“Does Ta have an inverse that leaves a named cost (fee, audit trail, ordering shift)?”
“Does Ta mutate any typed slot or external sink?” (no)
no annotation
calculate_* , get_* , list_* , search_* , find_*
Appendix
Table 16: Per-tool authoring decision table. Apply rows top-to-bottom; the first matching row determines the annotation. Read-only tools match no row and carry no Layer A entry. The recommended domain default policy is Veto when state-change tools dominate the action surface, Escalate when user-explicit-intent tools dominate.
Model
Domain
Pairs
Matched Δ
Task-cluster 95% CI
qwen3-8B
airline
143
+0.119
[+0.057,+0.184]
qwen3-8B
retail
150
−0.100
[−0.187,−0.020]
qwen3-8B
telecom
118
+0.034
[+0.000,+0.077]
llama-3.1-8b
airline
102
+0.137
[+0.049,+0.230]
llama-3.1-8b
retail
74
+0.000
[+0.000,+0.000]
llama-3.1-8b
telecom
70
+0.057
[+0.014,+0.115]
Appendix
Table 17: Matched task–trial reward differences. Each arm attempted 150 runs per cell; the intersection can be substantially smaller.