PACE: Provenance-Aware Capability Enforcement for Tool-Using LLM Agents
Organizations: PRADA Lab, King Abdullah University of Science and Technology · Imperfect Information Learning Team, RIKEN Center for Advanced Intelligence Project · State Key Laboratory of Internet of Things for Smart City, University of Macau · National Institute of Informatics · School of Computer Science and Engineering, University of Electronic Science and Technology of China
Abstract
Tool-using large language model (LLM) agents turn generated text into real side effects, so poisoned tool metadata, retrieved pages, memory, and reusable skills can steer the next call. Vetting an artifact before admission does not settle this. A safe variant and a leaking variant can produce the same admission evidence, and a sound gate then cannot relax that site for either. We make that condition precise, which leaves the last boundary a deployment can still act on. We present Provenance-Aware Capability Enforcement (PACE), which mediates every tool call immediately before it executes. Path confinement proposes an executable cut of represented influence paths, while capability and effect verification checks schema-defined effects against authority compiled from the authenticated request. We distinguish the certified execution contract from the evaluated configuration, which can restore an authorized call after a proposed block or apply a declared repair. Confinement requires the final action to preserve the certified cut. On eight executable agent-security benchmarks with three target-model families, the evaluated configuration gives strictly lowest attack success in 62 of 79 eligible attack columns and ties in 14; full-benchmark native utility loses at most three points relative to the undefended agent. A complete ablation over 1167 paired cases attributes most security gains to effect verification and refusal control to boundary adaptation. A reduced-scale adaptive search succeeds on 0/30 out-of-authority targets against the defense.
Figures & tables
| Qwen | DeepSeek | ChatGPT | ||||
| Method | ASR | Util. | ASR | Util. | ASR | Util. |
| AgentDojo | ||||||
| Baseline | 48.0 | 62.59 | 99.84 | 75.68 | 100.0 | 73.77 |
| CaMel | 0.32 | 35.14 | 28.14 | 60.04 | 36.09 | 57.02 |
| DTA | 0.95 | 29.84 | 33.33 | 54.0 | 55.01 | 56.01 |
| DataSent. | 35.29 | 73.77 | 78.06 | 85.0 | 84.1 | 82.95 |
| A0 | A1 | A2 | A3 | A4 | A5 | A6 | A7 | |
| Path confinement (P) | ✓ | ✓ | ✓ | ✓ | ||||
| Capability and effect (C) | ✓ | ✓ | ✓ | ✓ | ||||
| Boundary adaptation (B) | ✓ | ✓ | ✓ | ✓ | ||||
| Attack success (%), lower better | ||||||||
| WASP, plain text | 8.3 | 8.3 | 0.0 | 0.0 | 8.3 | 0.0 | 8.3 | 0.0 |
| WASP, URL injection | 8.3 | 8.3 | 8.3 | 0.0 | 0.0 | 0.0 | 16.7 | 0.0 |
Appendix figures & tables14 assets
Supplementary material from the paper’s appendix.
Appendix
| Notation | Description |
|---|---|
| Agent execution and threat model ( Section 2.1 ) | |
| Frozen policy that alternates free text with tool calls | |
| Registered tool set | |
| Tool call proposed at step | |
| Observation returned by | |
| Answer returned to the user, itself an egress site | |
| Field | Value |
|---|---|
| Target models | gpt-5.6-luna ; DeepSeek-V4-Flash ; local Qwen-3.8-27B served by vLLM |
| Decoding | Temperature ; maximum completion length tokens |
| Ablation seed | 20260828 ; single deterministic paired run |
| Concurrency | 8 workers; throughput only, no effect on the statistical population |
| Defense lanes | none , pace , and the external baselines of Table 7 |
| Layer switches | Path confinement, capability and effect verification, and execution-boundary adaptation, toggled independently for the ablation arms of Table 6 |
| Benchmark | Scaffold and native evaluator | Paired cases |
|---|---|---|
| AgentDojo | Official task, security, and utility functions | 198 |
| AgentDyn | Official three-domain runner and utility scorer | 60 |
| WASP | Official end-to-end evaluator; plain-text and URL groups | 24 |
| ASB | Official agents and metric scripts; seven run variants | 280 |
| PASB | Official personalized-agent workflow; IPI and four memory subsets | 58 |
| InjecAgent | Official tools and cases; Base and Enhanced splits | 212 |
| Arm | P | C | B | Role |
|---|---|---|---|---|
| A0 | Wrapper-matched no-defense corner | |||
| A1 | ✓ | Path confinement alone | ||
| A2 | ✓ | Capability and effect verification alone | ||
| A3 | ✓ | Execution-boundary adaptation alone | ||
| A4 | ✓ | ✓ | Path confinement with verification | |
| A5 | ✓ | ✓ | Path confinement with boundary adaptation |
| Family | Methods | Benchmarks and fidelity |
|---|---|---|
| No defense | Baseline, ReAct | All eight; native |
| Trusted flow and IFC | CaMeL ( Debenedetti et al., 2025 ) , FIDES ( Costa et al., 2025 ) , DRIFT ( Li et al., 2025a ) | AgentDojo, AgentDyn, WASP, PASB, MSB; native on AgentDojo, validated port elsewhere |
| Tool-use safeguards | Progent ( Shi et al., 2025 ) , Tool Filter ( Debenedetti et al., 2024 ) , Tool Allowlist ( Wang et al., 2026c ) , MCIP Guardian ( Jing et al., 2025 ) , ToolShield ( Li et al., 2026b ) | AgentDojo, ASB, AgentDyn, MCPTox, MSB, PASB; native or validated port |
| Re-execution | MELON ( Zhu et al., 2025 ) | AgentDojo; native |
| Detection | DataSentinel ( Liu et al., 2025 ) , PI-Detector ( Debenedetti et al., 2024 ) , PIGuard ( Li et al., 2025b ) , PromptGuard-2 ( Llama Team, 2025 ) , Metadata Sanitization ( Wang et al., 2026c ) | AgentDojo, InjecAgent, ASB, AgentDyn, MCPTox, PASB, MSB; pre-filter plus end-to-end |
| Prompt-level | Delimiters, Sandwich, Spotlighting, Instructional Prevention, Direct and PoT Paraphrase, PoT Shuffle, Repeat User Prompt | WASP, ASB, InjecAgent, AgentDyn, PASB; native |
| Method | Qwen-3.8-27B | DeepSeek-V4-Flash | gpt-5.6-luna | ||||||
|---|---|---|---|---|---|---|---|---|---|
| Plain-text | URL Injection | Utility | Plain-text | URL Injection | Utility | Plain-text | URL Injection | Utility | |
| ASR | ASR | Accuracy | ASR | ASR | Accuracy | ASR | ASR | Accuracy | |
| Baseline | 19.05 | 14.29 | 78.57 | 9.52 | 7.14 | 79.76 | 4.76 | 4.76 | 86.90 |
| Prompt Filter | 9.52 | 14.29 | 82.14 | 7.14 | 4.76 | 75.00 | 11.90 | 9.52 | 83.33 |
| Simple Static | 20.37 | 15.32 | 72.64 | 16.67 | 14.29 | 64.29 | 19.04 | 14.29 | 84.52 |
| FIDES | 2.38 | 40.48 | 0.0 | 19.05 | 30.95 | 32.14 | 4.76 | 23.81 | 59.52 |
| Method | Qwen-3.8-27B | DeepSeek-V4-Flash | gpt-5.6-luna | |||||||||
|---|---|---|---|---|---|---|---|---|---|---|---|---|
| Direct | Import Instructions | Tool Knowledge | Utility | Direct | Import Instructions | Tool Knowledge | Utility | Direct | Import Instructions | Tool Knowledge | Utility | |
| ASR | ASR | ASR | Accuracy | ASR | ASR | ASR | Accuracy | ASR | ASR | ASR | Accuracy | |
| Baseline | 48.00 | 30.21 | 30.52 | 62.59 | 97.62 | 99.68 | 99.84 | 75.68 | 99.05 | 99.84 | 100 | 73.77 |
| CaMeL | 0.32 | 0.32 | 0.32 | 35.14 | 20.51 | 24.48 | 28.14 | 60.04 | 31.00 | 36.09 | 34.02 | 57.02 |
| DTA | 0.48 | 0.95 | 0.95 | 29.84 | 24.17 | 28.46 | 33.33 | 54.00 | 42.13 | 48.01 | 55.01 | 56.01 |
| Spotlighting | 98.40 | 86.30 | 80.00 | 62.47 | 100 | 92.05 | 90.14 | 75.42 | 100 | 100 | 100 | 70.01 |
| Method | Qwen-3.8-27B | DeepSeek-V4-Flash | gpt-5.6-luna | |||||||||
|---|---|---|---|---|---|---|---|---|---|---|---|---|
| Direct Harm | DS-S1 | DS-S2 | Data-stealing | Direct Harm | DS-S1 | DS-S2 | Data-stealing | Direct Harm | DS-S1 | DS-S2 | Data-stealing | |
| ASR(%) | ASR(%) | ASR(%) | ASR(%) | ASR(%) | ASR(%) | ASR(%) | ASR(%) | ASR(%) | ASR(%) | ASR(%) | ASR(%) | |
| Baseline | 12.35 | 23.71 | 79.07 | 18.75 | 19.61 | 28.13 | 52.29 | 14.71 | 12.55 | 18.57 | 50.50 | 9.38 |
| ReAct | 24.31 | 27.94 | 92.11 | 25.74 | 27.84 | 37.68 | 79.02 | 29.77 | 16.86 | 23.35 | 66.14 | 15.44 |
| PromptGuard | 13.14 | 24.45 | 80.45 | 19.67 | 21.57 | 31.80 | 64.74 | 20.59 | 10.00 | 17.28 | 56.38 | 9.74 |
| Prompt-Sandwich | 0.59 | 0.18 | 0.15 | 0.12 | 4.71 | 14.89 | 50.62 | 7.54 | 2.16 | 17.10 | 73.12 | 12.50 |
| Method | Qwen-3.8-27B | DeepSeek-V4-Flash | gpt-5.6-luna | |||||||||||||||
|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|
| DPI | OPI | MP | Mixed | PoT | Utility | DPI | OPI | MP | Mixed | PoT | Utility | DPI | OPI | MP | Mixed | PoT | Utility | |
| ASR(%) | ASR(%) | ASR(%) | ASR(%) | ASR(%) | Accuracy(%) | ASR(%) | ASR(%) | ASR(%) | ASR(%) | ASR(%) | Accuracy(%) | ASR(%) | ASR(%) | ASR(%) | ASR(%) | ASR(%) | Accuracy(%) | |
| Baseline | 84.92 | 66.00 | 13.25 | 88.50 | 56.25 | 30.29 | 74.08 | 58.25 | 36.50 | 70.75 | 26.00 | 64.00 | 81.08 | 31.50 | 27.00 | 72.75 | 73.50 | 80.25 |
| Delimiters | 88.08 | 73.75 | 13.50 | 86.75 | 23.75 | 24.61 | 88.42 | 85.50 | 51.75 | 93.00 | 38.00 | 54.39 | 81.63 | 68.25 | 48.25 | 88.00 | 76.25 | 66.57 |
| Sandwich | 85.33 | 61.75 | 11.75 | 88.50 | 68.75 | 32.50 | 81.83 | 76.25 | 46.00 | 59.75 | 25.75 | 69.07 | 71.21 | 61.50 | 52.00 | 63.25 | 57.25 | 87.86 |
| Instructional Prevention | 71.92 | 60.75 | 17.00 | 82.00 | 52.75 | 33.93 | 58.75 | 61.75 | 40.75 | 78.50 | 20.75 | 72.50 | 54.57 | 40.75 | 43.25 | 83.75 | 46.00 | 87.25 |
| Method | Qwen-3.8-27B | DeepSeek-V4-Flash | gpt-5.6-luna | |||||||||
|---|---|---|---|---|---|---|---|---|---|---|---|---|
| Shopping | Github | DailyLife | Utility | Shopping | Github | DailyLife | Utility | Shopping | Github | DailyLife | Utility | |
| ASR(%) | ASR(%) | ASR(%) | Utility(%) | ASR(%) | ASR(%) | ASR(%) | Utility(%) | ASR(%) | ASR(%) | ASR(%) | Utility(%) | |
| Baseline | 11.11 | 17.22 | 32.50 | 63.39 | 12.78 | 17.22 | 19.50 | 65.00 | 7.78 | 10.56 | 19.50 | 74.82 |
| Repeat User Prompt | 6.67 | 15.56 | 34.50 | 61.25 | 22.78 | 10.00 | 31.50 | 61.43 | 12.28 | 10.00 | 31.50 | 70.54 |
| Spotlighting | 3.33 | 13.89 | 5.50 | 65.00 | 13.33 | 7.22 | 3.00 | 58.21 | 17.22 | 15.00 | 27.50 | 64.46 |
| Tool Filter | 0.0 | 0.0 | 0.05 | 7.86 | 1.67 | 2.78 | 2.00 | 8.04 | 2.78 | 3.89 | 6.50 | 9.11 |
| Method | Qwen-3.8-27B | DeepSeek-V4-Flash | gpt-5.6-luna | |||||||||||||||
|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|
| IPI | Short Read | Short Modify | Long Read | Long Modify | Utility | IPI | Short Read | Short Modify | Long Read | Long Modify | Utility | IPI | Short Read | Short Modify | Long Read | Long Modify | Utility | |
| ASR(%) | ASR(%) | ASR(%) | ASR(%) | ASR(%) | Accuracy(%) | ASR(%) | ASR(%) | ASR(%) | ASR(%) | ASR(%) | Accuracy(%) | ASR(%) | ASR(%) | ASR(%) | ASR(%) | ASR(%) | Accuracy(%) | |
| Baseline | 0.0 | 85.00 | 62.50 | 62.50 | 75.00 | 98.47 | 0.0 | 67.50 | 72.50 | 60.00 | 92.50 | 99.24 | 2.29 | 55.00 | 42.50 | 50.00 | 97.5 | 100.0 |
| Delimiters | 0.0 | 25.00 | 0.0 | 45.00 | 20.00 | 99.23 | 0.0 | 30.00 | 55.00 | 42.50 | 32.50 | 98.47 | 1.53 | 22.50 | 35.00 | 30.00 | 35.00 | 99.24 |
| Sandwich | 0.0 | 85.00 | 70.00 | 65.00 | 72.50 | 99.24 | 0.0 | 72.50 | 85.00 | 57.50 | 87.50 | 98.47 | 3.82 | 40.00 | 25.00 | 42.50 | 82.50 | 100.0 |
| Instructional Prevention | 0.0 | 17.50 | 0.0 | 17.50 | 0.0 | 97.67 | 0.0 | 20.00 | 17.50 | 27.50 | 17.50 | 96.95 | 6.87 | 17.50 | 7.50 | 27.50 | 25.00 | 98.47 |
| Method | Qwen-3.8-27B | DeepSeek-V4-Flash | gpt-5.6-luna | |||||||||
|---|---|---|---|---|---|---|---|---|---|---|---|---|
| Template-1 | Template-2 | Template-3 | Refusal | Template-1 | Template-2 | Template-3 | Refusal | Template-1 | Template-2 | Template-3 | Refusal | |
| ASR(%) | ASR(%) | ASR(%) | Refusal Rate(%) | ASR(%) | ASR(%) | ASR(%) | Refusal Rate(%) | ASR(%) | ASR(%) | ASR(%) | Refusal Rate(%) | |
| Baseline | 22.39 | 8.84 | 32.52 | 0.16 | 56.54 | 52.61 | 56.06 | 0.42 | 23.04 | 26.10 | 86.93 | 27.96 |
| Metadata Sanitization | 5.94 | 1.38 | 17.41 | 0.15 | 13.09 | 9.91 | 28.60 | 0.37 | 5.76 | 5.43 | 91.48 | 26.67 |
| PI-Detector | 20.53 | 7.37 | 25.22 | 0.40 | 85.34 | 4.59 | 50.19 | 1.63 | 48.17 | 3.97 | 95.45 | 30.14 |
| PI-Guard | 11.59 | 3.09 | 11.07 | 0.91 | 50.79 | 27.56 | 30.68 | 3.15 | 34.03 | 9.19 | 63.64 | 46.32 |
| Method | Qwen-3.8-27B | DeepSeek-V4-Flash | gpt-5.6-luna | |||||||||||||||
|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|
| Plan | Call | Response | Multi-Stage | PUA | NRP | Plan | Call | Response | Multi-Stage | PUA | NRP | Plan | Call | Response | Multi-Stage | PUA | NRP | |
| ASR(%) | ASR(%) | ASR(%) | ASR(%) | Accuracy(%) | Accuracy(%) | ASR(%) | ASR(%) | ASR(%) | ASR(%) | Accuracy(%) | Accuracy(%) | ASR(%) | ASR(%) | ASR(%) | ASR(%) | Accuracy(%) | Accuracy(%) | |
| Baseline | 5.67 | 38.75 | 26.35 | 18.70 | 46.34 | 36.29 | 61.33 | 100.0 | 29.68 | 52.00 | 91.24 | 49.29 | 67.00 | 98.75 | 9.03 | 50.20 | 89.08 | 54.51 |
| MCIP Guardian | 13.33 | 50.00 | 14.03 | 8.40 | 59.15 | 55.70 | 8.00 | 77.50 | 26.13 | 37.30 | 96.55 | 51.32 | 29.33 | 76.25 | 8.55 | 36.90 | 92.34 | 56.87 |
| PI-Detector | 16.67 | 38.75 | 7.02 | 5.20 | 54.36 | 49.83 | 55.33 | 53.75 | 18.55 | 48.40 | 94.81 | 51.18 | 65.00 | 70.00 | 7.10 | 47.10 | 90.85 | 56.22 |
| PI-Guard | 14.00 | 37.50 | 13.87 | 9.00 | 56.42 | 53.26 | 32.00 | 57.50 | 36.94 | 36.90 | 91.98 | 49.95 | 40.67 | 66.25 | 10.00 | 38.50 | 89.97 | 54.25 |
| Quantity | Statistic | Interpretation |
|---|---|---|
| forward_simulation_ observed=true | Rate + CI | Last event required no graph expansion; not proof of H1 |
| op.unknown expansions | Histogram/quantiles | Observed abstraction incompleteness |
| Unbindable minimum cuts | Rate + CI | Graph separation lacking an executor action |
| Stale-state rejection | Rate + CI | Freshness check exercised before dispatch |
| Effect-obligation failures | Histogram by obligation | Which of the nine checks fails, and how often |
| Bound action executed | Rate + CI | Observable executor agreement relevant to H4 |