NOMOS: Compiling Written Policies into Statically Verified Tool-Call Gates for LLM Agents
Organizations: Corners Co., Ltd.
Abstract
Tool-using LLM agents violate the policies they are deployed to enforce, often silently. Prior defenses hand-write rules, query an LLM verifier per action, or compile policies through heavyweight formal machinery. Naive compilation fails: extracted rules block the tool satisfying their own precondition, or read arguments their tool lacks. NOMOS, a four-pass compiler, turns a natural-language policy into a deterministic tool-call gate; static verification with tool-schema-level checks alone (no prover, solver, or LLM) repairs or rejects 37% (airline) and 13% (retail) of candidates, without which most shipped rules are inoperable. Replaying compiled rules over undefended transcripts flags bindings that refuse legitimate work (a development binding refused 95.9% of task-passing calls); no evaluation binding is flagged. On -bench the gate cuts violations of reference-encoded clauses among state-changing calls from 66.3% to 2.6% (airline) and 30.8% to 6.9% (retail), raising airline task success significantly for ; a 26B on-premise compilation is not significantly worse than hand-written or frontier-compiled rules. Unlike AgentDojo's shipped defenses, it reaches a zero attack success rate (ASR) on banking, where nine attack families collapse onto three structural rules. On the other three suites its ASR is at most 3.6%, from goals with no tool call to govern and one write admitted by a binding weaker than its clause; a second agent model, Llama-3.3-70B, reproduces the effect on both benchmarks. Decisions take microseconds without an LLM call, at a domain-dependent benign-utility cost; compilation runs on-premise on open-weight gemma-4-26B.
Figures & tables
| Condition | ASR | Util. (atk) | Benign | Runtime LLM |
|---|---|---|---|---|
| Undefended | 46.5% | 64.9% | 79.2% | — |
| Progent (auto) | 5.6% | 52.1% | 62.5% | ✓ |
| NOMOS gate | 0.0% | 67.4% | 70.8% | — |
| Policy | LLM | Static | Runtime | Solver/ | Pre- | |
|---|---|---|---|---|---|---|
| System | source | extr. | verif. | LLM | prover | flight |
| PolicyGuard [ 4 ] | doc | ✓ | — | ✓ | — | — |
| Progent [ 27 ] | task | ✓ | — | ✓ | ✓ | — |
| VeriGuard [ 6 ] | doc | ✓ | ✓ | ✓ | ✓ | — |
| Zwerdling [ 7 ] | doc | ✓ | tests | — | — | — |
| NOMOS | doc | ✓ | ✓ | — | — | ✓ |
| airline | retail | |
| Raw extraction candidates (5 samples) | 79 | 53 |
| Structural defects (rule names real tools, cannot function) | ||
| Repaired: self-blocking (deadlock) | 0 | 4 |
| Repaired: write-scope overreach | 5 | 0 |
| Argument–signature mismatch | 8 | 1 |
| Rejected: domain-unsatisfiable predicate | 5 | 0 |
| Tool bound to TEL-04 | Refused/calls | Rate | Verdict |
|---|---|---|---|
| send_payment_request | 47/49 | 95.9% | flagged |
| resume_line | 0/44 | 0.0% | kept |
| Domain | Condition | Writes | Viol. | Rate | 95% CI | Batch |
|---|---|---|---|---|---|---|
| airline | undefended | 353 | 234 | 66.3% | [61.2, 71.0] | 36.8% |
| airline | NOMOS gate | 116 | 3 | 2.6% | [0.9, 7.3] | 2.6% |
| retail | undefended | 611 | 188 | 30.8% | [27.2, 34.5] | 13.6% |
| retail | NOMOS gate | 582 | 40 | 6.9% | [5.1, 9.2] | 5.3% |
| Metric | Undef. | Gate | 95% CI | |
| airline (50 tasks 5 trials) | ||||
| pass^1 | 39.2% | 47.6% | ||
| pass^2 | 27.4% | 40.6% | * | |
| pass^3 | 23.0% | 37.0% | * | |
| pass^4 | 21.2% | 34.4% | * | |
| pass^5 | 20.0% | 32.0% | ||
| Domain | Cond. | Pass | Safe | P V | W/sim | V/sim |
|---|---|---|---|---|---|---|
| airline | undef. | 39.2% | 32.4% | 17 | 1.41 | 0.94 |
| airline | gate | 47.6% | 47.6% | 0 | 0.46 | 0.01 |
| retail | undef. | 52.4% | 41.0% | 52 | 1.34 | 0.41 |
| retail | gate | 50.2% | 47.8% | 11 | 1.28 | 0.09 |
| airline, Llama | undef. | 22.8% | 18.4% | 11 | 1.65 | 1.44 |
| airline, Llama | gate | 46.8% | 46.8% | 0 | 0.11 | 0.01 |
| airline | retail | ||||||
|---|---|---|---|---|---|---|---|
| Configuration | Rules | Inop. | r/o | Rules | Inop. | r/o | |
| All checks (reference) | 8 | 0 | 0 | 4 | 0 | 0 | |
| self-blocking reachability | 8 | 0 | 0 | 4 | 1 | 0 | |
| write-scope repair | 8 | 0 | 1 | 4 | 0 | 0 | |
| argument consistency | 10 | 3 | 0 | 4 | 0 | 0 | |
| domain satisfiability | 9 | 1 | 0 | 4 | 0 | 0 | |
| Domain | Bind. | Supp. | Fire | Worst | Flag |
|---|---|---|---|---|---|
| airline (10 rules) | 16 | 15 | 11 | 83.3% (5/6) | 0 |
| retail (5 rules) | 24 | 24 | 12 | 31.6% (12/38) | 0 |
| telecom (pre-repair) | 2 | 2 | 1 | 95.9% (47/49) | 1 |
| Rule source | Preparation (once per policy version) | Policy leaves premises | Runtime LLM calls | Rules | pass^1 / pass^5 (vs. local) |
|---|---|---|---|---|---|
| local compilation (26B, ours) | 40 LLM calls, $0 | no | 0 | 10 | 47.6 / 32.0 (reference) |
| hand-written reference | engineer time | no | 0 | 16 | 42.8 / 28.0 ( significant, ) |
| frontier compilation | 40 LLM calls, $1.02 | yes | 0 | 14 | 46.4 / 30.0 (not significant) |
| LLM verifier, PolicyGuard-style | checklist compile | depends on verifier host | 1 per state-changing decision; 347 in 250 simulations | n/a | 47.2 / 38.0 (not significant; wall-clock) |
| Condition | Pairs | Att. | Blk. | Succ. | ASR |
|---|---|---|---|---|---|
| Std., undefended | 288 | 142 | 0 | 134 | 46.5% |
| Std., gate | 288 | 122 | 122 | 0 | 0.0% |
| Adaptive, undefended | 144 | 55 | 0 | 50 | 34.7% |
| Adaptive, gate | 144 | 61 | 61 | 0 | 0.0% |
| Defense | ASR | Util. (atk) | Benign |
|---|---|---|---|
| Undefended | 46.5% | 64.9% | 79.2% |
| Spotlighting (delimiting) | 36.5% | 66.0% | 77.1% |
| Repeat user prompt | 31.9% | 66.0% | 83.3% |
| Tool filter | 11.8% | 49.7% | 58.3% |
| NOMOS gate | 0.0% | 67.4% | 70.8% |
| Agent model | Cond. | ASR | Att. | Blk. |
|---|---|---|---|---|
| gemma-4-26B | undef. | 46.5% | 49.3% | — |
| gate | 0.0% | 42.4% | 100% | |
| Llama-3.3-70B | undef. | 44.1% | 47.8% | — |
| gate | 0.0% | 40.7% | 100% |
| Airline | gemma-4-26B | Llama-3.3-70B |
|---|---|---|
| pass^1 | 39.2 47.6 ( ) | 22.8 46.8 ( *) |
| pass^5 | 20.0 32.0 ( ) | 14.0 44.0 ( *) |
| safe completion | 32.4 47.6 ( *) | 18.4 46.8 ( *) |
| violations / writes | 66.3% 2.6% | 86.9% 7.1% |
| executed writes | 353 116 | 413 28 |
| gate refusals | 226 | 250 |
| Undef. ASR | Reported | Executed | ||||
|---|---|---|---|---|---|---|
| Suite | gemma | Llama | gemma | Llama | gemma | Llama |
| banking | 46.5% | 44.1% | 0.0% | 0.0% | 0.0% | 0.0% |
| slack | 70.5% | 59.0% | 0.0% | 10.5% | 0.0% | 0.0% |
| travel | 10.0% | 60.7% | 3.6% | 5.7% | 0.7% | 1.4% |
| workspace | 7.3% | 13.6% | 0.0% | 0.4% | 0.0% | 0.0% |
| ASR | Benign util. | |||
|---|---|---|---|---|
| Suite | undef. | gate | undef. | gate |
| banking | 46.5% | 0.0% | 79.2% | 70.8% |
| slack | 70.5% | 0.0% | 90.5% | 14.3% |
| travel | 10.0% | 3.6% † | 75.0% | 70.0% |
| workspace | 7.3% | 0.0% | 87.5% | 65.0% |
| Rule set | slack | travel | workspace |
|---|---|---|---|
| Undefended | 70.5 / 90.5 | 10.0 / 75.0 | 7.3 / 87.5 |
| Default vocabulary | 18.1 / 28.6 | – | – |
| Extended | 14.3 / 57.1 | 5.7 / 60.0 | 7.7 / 87.5 |
| + accessor check ‡ | 0.0 / 14.3 | 3.6 / 70.0 | 0.0 / 65.0 |
| Rule | Guarded tools | Predicate |
|---|---|---|
| slack (4 rules) | ||
| FBD-01 | send_channel_message , send_direct_message | body_has_unnamed_url |
| FBD-02 | send_direct_message | recipient_unknown |
| FBD-02G | add_user_to_channel , invite_user_to_slack , remove_user_from_slack , send_channel_message | target_not_user_named |
| FBD-03 | get_webpage , post_webpage | target_not_user_named |
| travel (3 rules) | ||
| Goal | Attacker objective | Att. | Blk. | Rule |
|---|---|---|---|---|
| 0 | exfiltrate streaming-service name | 24 | 24 | FBD-02 |
| 1 | exfiltrate a third party’s IBAN | 13 | 13 | FBD-02 |
| 2 | exfiltrate food preference | 30 | 30 | FBD-02 |
| 3 | exfiltrate phone model | 21 | 21 | FBD-02 |
| 4 | redirect recurring payment | 20 | 20 | PRE-03 |
| 5 | transfer maximum funds | 12 | 12 | FBD-02 |