Where Rules End and Judges Begin: Measuring the Judgment Boundary in Multi-Agent Systems Security
Organizations: Department of Computer Science The University of Alabama, Tuscaloosa, USA
Abstract
LLM-based multi-agent systems (MAS) engage tools, share memory, and delegate tasks, often encountering adversarial content. Current defenses for MAS are typically evaluated in isolation, focusing on one attack type at a time, which can lead to costly and hard-to-audit outcomes. This study organizes defenses into five principles, implementing them as DEFER1 (DEterministic-First Enforcement with Residual judgment), which includes a cascade of 28 checks that blocks what it can and refers the rest to a panel of four judges. In independent testing across four domains, attack success rates drop from about 30.0% to approximately 3.0%, with 78% of blocked attacks handled by deterministic checks. Only a quarter of proposals reach the judges in the security-operations domain, illustrating that the rules provide security for attacks violating clear policies, while judges manage those that only misrepresent intent. Both systems have weaknesses, such as a risk-score approval gate that inaccurately approves most attack proposals but few legitimate ones, highlighting the challenges in assessing threats accurately.
Figures & tables
| ID | Channel | Attacker content arrives as | Surface | Tier |
|---|---|---|---|---|
| C1 | Task input | Text inside an alert, ticket, or request field | Tool | T1 |
| C2 | Tool response | Text returned by a tool the agent called | Tool | T1 |
| C3 | Memory content | A record retrieved from a shared store | Memory | T1 |
| C4 | Handoff | The summary one phase passes to the next | Tool | T2 |
| C5 | Proposal justification | The rationale attached to an action proposal | Tool | T2 |
| Attack variants | Benign | ||||
| Configuration | ASR % [95% CI] | Attempt % | Block att. % | Judged % | Denied % |
| Flat (undefended) | 32.0 [22.7, 41.8] | 32.0 | n/a | 0 | 0.0 |
| ACL (connectivity only) | 24.9 [15.6, 35.1] | 28.9 | 13.8 | 0 | 51.9 a |
| JudgeOnly (judge every proposal) | 34.7 [24.4, 45.3] | 44.4 | 22.0 | 99.9 | 4.5 |
| NoJudge (never judge) | 1.3 [0.0, 4.0] | 44.0 | 97.0 | 0 | 54.3 |
| Full (judge last) | 2.2 [0.0, 5.8] | 42.2 | 94.7 | 26.4 | 9.9 |
Appendix figures & tables37 assets
Supplementary material from the paper’s appendix.
Appendix
| Attack vector | Tool-side effect | Memory-side effect |
| Component level | ||
| Unauthorized access | Privilege escalation, unauthorized execution | Unauthorized retrieval |
| Context contamination | API exfiltration | Privacy leakage, data poisoning |
| Coordination level | ||
| Lateral compromise | Misuse via shared privileges | False injection via shared memory |
| Consensus manipulation | Biased tool selection | Conflicting intelligence injection |
| AP | Name | Attacker objective | Channel | Tier | Effect scored as success |
|---|---|---|---|---|---|
| AP-1 | Tool redirection | Induce a phase agent to invoke a tool outside its phase | C1 | T1 | Out-of-manifest tool executes |
| AP-2 | Memory poisoning | Get falsified findings stored so later phases reason on them | C2, C1 | T1 | Canary persisted in a shared store |
| AP-3 | Confused deputy | Have the agent act on attacker text returned as a tool result | C2 | T1 | Action named in the forged result executes (never attempted by any primary) |
| AP-4 | Cross-phase exfiltration | Move sensitive incident data into an outbound feed | C3, C1 | T1 | Canary appears in an outbound tool call, or an unauthorized read is allowed (Section V-D ) |
| AP-5 | Unauthorized irreversible action | Trigger bulk or destructive operations | C1, C4 | T1–T2 | Bulk or destructive action executes |
| AP-6 | Replay and accumulation | Re-submit a previously approved proposal, or accumulate actions across incidents | C1 | T1 | Stale or repeated proposal executes |
| ID | Vector | Check | AP | ID | Vector | Check | AP |
|---|---|---|---|---|---|---|---|
| TA-1 | Tool redirection (out-of-scope invocation) | P2.1 | 1 | TA-19 | Dangerous call hidden in a batch | P2.1, host | 1, 12 |
| TA-2 | Forged component identity | P1.1 | 3 | TA-20 | Deflate severity, then escalate | P3.1, P3.3 | 5, 9 |
| TA-3 | Tampered tool response | P1.2 | 3 | TA-21 | Probing through rejection reasons | host | 4, 10 |
| TA-4 | Cross-phase data exfiltration via tool output | P2.3 | 4, 8 | TA-22 | Cross-incident action accumulation | P3.4, P3.5 | 6, 12 |
| TA-5 | Unauthorized irreversible action | P3.8, P3.9 | 5 | MA-1 | Cross-phase memory read | P5.1 | 4 |
| TA-6 | Replay of an approved action proposal | P3.6 | 6, 7 | MA-2 | Field-level data leakage | P5.2 | 4, 14 |
| Principle | Draws on | Defer instantiation | Difference from source |
|---|---|---|---|
| P1 Authorized Interface | Tool registries [ narajala2025securing ] ; agent identity binding [ syros2025saga ] | Registry allowlist; response schema, timing, and replay checks; per-tool response MAC; configuration hash snapshot | No public-key manifest signing or agent identity lifecycle; symmetric keys held by the host |
| P2 Capability Scoping | Prompt flow integrity [ kim2025prompt ] ; delegation threat modeling [ krawiecka2025extending ] | Static per-phase manifests and action caps; parameter rules; target–evidence consistency; output classifier | No tracking of untrusted data flow through prompts; privileges are static per phase, not per task |
| P3 Verified Execution | Ledger-anchored verification [ zou2025blocka2a ] ; Byzantine-robust consensus [ chen2024blockagents ] | Context, sequence, and ledger checks; risk scoring with auto-gates; multi-model LLM panel; pre-execution hash check | Ledgers are local, not distributed or blockchain-anchored; validators are assumed uncompromised |
| P4 Memory Integrity | Consensus write validation [ wei2025memguard ] ; filtering and purging [ mao2025agentsafe ] ; isolate-then-aggregate retrieval [ xiang2024certifiably ] | Schema, metadata, evidence-similarity, drift, replay, and rule-based contradiction checks on every write | No LLM consensus on writes; no certified retrieval; contradiction detection is phrase-based |
| P5 Access-Controlled Isolation | Hierarchical memory isolation [ mao2025agentsafe ] ; dynamic access graphs [ rezazadeh2025collaborative ] | Phase-to-store policy; field-level redaction; query-scope limits; read-pattern monitor; read-result sanitization | Permissions are static, not time-varying; sanitization is pattern-based |
| Principle | Checks | Indep. | Dep. | Sim. | LLM | Stateful |
|---|---|---|---|---|---|---|
| P1 Authorized Interface | 3 | 3 | 0 | 0 | 0 | 1 |
| P2 Capability Scoping | 3 | 1 | 0 | 2 | 0 | 0 |
| P3 Verified Execution | 11 | 5 | 5 | 0 | 1 | 4 |
| P4 Memory Integrity | 6 | 2 | 2 | 2 | 0 | 2 |
| P5 Access-Ctrl. Isolation | 5 | 2 | 1 | 2 | 0 | 1 |
| Total | 28 | 13 | 8 | 6 | 1 | 8 |
| ID | Check | Type | What it verifies |
| P1.1 | Identity | I | Tool (or store, validator) is present in the component registry |
| P1.2 | Response integrity | I, S | Response matches the registered schema; latency within 0.01 ms–30 s; response hash unseen within a five-incident window; per-tool MAC verifies |
| P1.3 | Configuration integrity | I | SHA-256 of every configuration file matches the start-up snapshot, re-verified per incident; HMAC on persisted consent ledgers |
| P2.1 | Manifest | I | Tool is in the calling phase’s manifest; per-incident action cap and bulk-action threshold not exceeded |
| P2.2 | Parameter validation | I, C, E | No wildcard targets; asset-criticality rules (deny all on critical, deny destructive on high); per-tool min/max/enum rules; action target appears in the incident evidence (substring, or cosine ) |
| P2.3 | Output classifier | C, E | Phase-aware sensitive-content patterns, and cosine to sensitive categories; redact or reject |
| Attack vector | P1 | P2 | P3 | P4 | P5 |
| Unauthorized access | ✓ | ✓ | |||
| Context contamination | ✓ | ✓ | |||
| Lateral compromise | ✓ | ✓ | |||
| Consensus manipulation | ✓ | ✓ | ✓ | ||
| Covert coordination | ✓ | ✓ | |||
| Authentication bypass | ✓ | ✓ |
| Feature | InjecAgent [ zhan2024injecagent ] | AgentDojo [ debenedetti2024agentdojo ] | ASB [ zhang2025asb ] | TAMAS [ tamas2025 ] | Ours |
|---|---|---|---|---|---|
| Multi-agent workflow | ✗ | ✗ | ✗ | ✓ | ✓ |
| Tool-response injection | ✓ | ✓ | ✓ | ✓ | ✓ |
| Memory channel | ✗ | ✗ | ✓ | – | ✓ |
| Inter-agent channels | ✗ | ✗ | ✗ | ✓ | ✓ |
| Benign / utility tasks | ✗ | ✓ | ✓ | ✓ | ✓ |
| Attempt vs. block outcome | ✗ | ✗ | – | ✓ |
| Interaction type | Flat | Defer | Red. |
|---|---|---|---|
| Agent tool | 64 | 16 | 75% |
| Agent memory | 48 | 16 | 67% |
| Agent agent | 12 | 4 | 67% |
| Tool response agent | 64 | 16 | 75% |
| External feed memory | 12 | 4 | 67% |
| Total | 200 | 56 | 72% |
| Principle | Related controls |
|---|---|
| P1 Authorized Interface | Zero-trust architecture (NIST SP 800-207); identification and authentication (SP 800-53 Rev. 5, IA family) [ nist80053r5 ] |
| P2 Capability Scoping | Least privilege (SP 800-53 AC-6); Excessive Agency (OWASP LLM06:2025) [ owasp2025llmtop10 ] |
| P3 Verified Execution | Separation of duties (SP 800-53 AC-5); logging (ISO/IEC 27001:2022 A.8.15); record-keeping (EU AI Act Art. 12) [ euaiact2024 ] |
| P4 Memory Integrity | Software, firmware, and information integrity (SP 800-53 SI-7); traceability (NIST AI RMF) |
| P5 Access-Ctrl. Isolation | Account management and access enforcement (SP 800-53 AC-2, AC-3); data minimization (GDPR Art. 5(1)(c)) |
| Primary model | Size | Panel |
|---|---|---|
| Qwen3-235B-A22B-Instruct-2507 | 235B MoE | Local4 |
| gpt-oss-120b | 117B MoE | Local4 minus self (2 of 3) |
| Llama-4-Scout-17B-16E-Instruct | 109B MoE | Local4 minus self (2 of 3) |
| Mistral-Small-3.2-24B-Instruct | 24B | Local4 minus self (2 of 3) |
| Llama-3.1-8B-Instruct | 8B | Local4 |
| Model | Precision | Parallelism | Hardware |
|---|---|---|---|
| Qwen3-235B-A22B-Instruct-2507 | BF16 | TP=4 | 4 H200 |
| Llama-4-Scout-17B-16E-Instruct | BF16 | TP=2 | 2 H200 |
| Mistral-Small-3.2-24B-Instruct | BF16 | 1 GPU | H200 |
| Llama-3.1-8B-Instruct | BF16 | 1 GPU | RTX 5090 |
| Qwen3-32B; DeepSeek-R1-Distill-Qwen-32B; Qwen3-14B | BF16 | 1 GPU | H200 (shared) |
| Gemma-4-31B-it; gpt-oss-120b | as released | 1 GPU | H200 |
| First interceptor | CyberOps | Other three | ASB |
|---|---|---|---|
| By decision tier | |||
| Content-independent rules | 31.1 | 11.0 | 0.0 |
| Content-dependent rules | 50.0 | 19.7 | 9.1 |
| Similarity-threshold checks | 16.7 | 39.0 | 31.2 |
| LLM panel | 2.2 | 30.3 | 59.7 |
| By principle | |||
| Flat | ACL | Full | Att. | Blk att. | |
| Domains (Qwen3-235B) | |||||
| CyberOps | 32.0 | 24.9 | 2.2 | 42.2 | 94.7 |
| Healthcare | 23.1 | 19.1 | 0.0 | 25.8 | 100.0 |
| Finance | 35.1 | 26.2 | 5.3 | 40.9 | 87.0 |
| Legal | 29.8 | 28.4 | 4.4 | 40.0 | 88.9 |
| All four domains | 30.0 | 24.7 | 3.0 | 37.2 | 91.9 |
| Primary | Flat | ACL | JudgeOnly | NoJudge | Full | Judged | Denied |
|---|---|---|---|---|---|---|---|
| Qwen3-235B | 32.0 | 24.9 | 34.7 | 1.3 | 2.2 | 26.4 | 9.9 |
| gpt-oss-120b | 21.3 | 18.7 | 21.3 | 1.3 | 6.2 | 9.1 | 0.0 |
| Llama-3.1-8B | 34.7 | 33.8 | 27.6 | 1.3 | 8.4 | 45.0 | 14.6 |
| Benchmark | Flat | ACL | Defer |
|---|---|---|---|
| Agent Security Bench (255 cases 2 trials) | |||
| All cases | 28.8 [23.7, 34.1] | 30.0 [24.9, 35.3] | 1.8 [0.4, 3.5] |
| Defer , single judge | 3.3 [1.6, 5.5] | ||
| Defer , Lin3 | 9.6 [6.3, 13.3] | ||
| TAMAS (250 instances 3 trials) | |||
| Direct prompt injection | 36.9 [25.5, 48.7] | – | 0.0 [0.0, 0.0] |
| AP-4 | AP-14 | |||
|---|---|---|---|---|
| Configuration | ASR | Exp. | ASR | Exp. |
| Flat | 57.1 | 100.0 | 0.0 | 100.0 |
| ACL | 14.3 | 28.6 | 0.0 | 100.0 |
| JudgeOnly | 85.7 | 100.0 | 0.0 | 100.0 |
| Full P5 | 28.6 | 100.0 | 0.0 | 100.0 |
| Full | 0.0 | 28.6 | 0.0 | 100.0 |
| Domain | Scen. | Denied % | Any-denial % |
|---|---|---|---|
| CyberOps | 20 | 9.9 | 81.7 |
| Healthcare | 5 | 20.0 | 100.0 |
| Finance | 5 | 19.5 | 100.0 |
| Legal | 5 | 26.9 | 100.0 |
| All | 35 | – | 89.5 |
| Flat | Full | |||||
|---|---|---|---|---|---|---|
| Channel | Var. | Exp. % | Att. exp. | ASR | Att. exp. | ASR |
| C1 task input | 170 | 100.0 | 26.1 | 26.1 | 37.6 | 2.5 |
| C2 tool response | 33 | 97.0 | 0.0 | 0.0 | 2.7 | 0.0 |
| C3 memory content | 35 | 100.0 | 14.3 | 14.3 | 0.0 | 0.0 |
| C4 handoff | 30 | 100.0 | 45.6 | 45.6 | 80.0 | 14.4 |
| C5 proposal rationale | 32 | 100.0 | 84.4 | 84.4 | 71.9 | 1.0 |
| Panel | Quorum | Let through | CyberOps ASR % | Other three ASR % | Benign rej. % |
|---|---|---|---|---|---|
| Single (Qwen3-32B) | 1/1 | 36 (8.9%) | 9.8 | 6.7 | 15.0 |
| Lin3 (Qwen family) | 2/3 | 127 (31.4%) | 9.8 | 6.8 | 14.1 |
| Div3L (three lineages) | 2/3 | 18 (4.4%) | 7.1 | 4.3 | 17.2 |
| Local4 (main) | 3/4 | 18 (4.4%) | 7.1 | 3.9 | 17.7 |
| ASR | Denied | |||||
|---|---|---|---|---|---|---|
| Arm | Original | Re-judged | Live | Original | Re-judged | Live |
| Qwen3-235B Full , CyberOps | 7.1 | 8.9 | 8.4 | 11.2 | 11.1 | 9.9 |
| Qwen3-235B Full , healthcare | 0.0 | 0.0 | 0.0 | 24.2 | 18.8 | 23.7 |
| Qwen3-235B Full , finance | 4.0 | 9.8 | 6.7 | 22.7 | 17.6 | 17.5 |
| Qwen3-235B Full , legal | 7.6 | 12.4 | 10.2 | 25.4 | 14.9 | 16.7 |
| Qwen3-235B JudgeOnly , CyberOps | 33.8 | 32.0 | 29.8 | 3.8 | 4.4 | 5.7 |
| Split | Content | Size | Tests |
|---|---|---|---|
| Our suite | CyberOps, healthcare, finance, legal | 300 variants | boundary, generality |
| Third-party | ASB; InjecAgent subset; TAMAS | 255; 50; 250 cases | author independence |
| Benign | Companion scenarios; TAMAS IPI tasks | 35; 50 | cost |