PACE: Provenance-Aware Capability Enforcement for Tool-Using LLM Agents
Authors: Fengpeng Li, Qizhou Wang, Yuke Hu, Kemou Li, Jun Liu, Haiwei Wu, Jiantao Zhou, Di Wang
Organizations: PRADA Lab, King Abdullah University of Science and Technology · Imperfect Information Learning Team, RIKEN Center for Advanced Intelligence Project · State Key Laboratory of Internet of Things for Smart City, University of Macau · National Institute of Informatics · School of Computer Science and Engineering, University of Electronic Science and Technology of China
Tool-using large language model (LLM) agents turn generated text into real side effects, so poisoned tool metadata, retrieved pages, memory, and reusable skills can steer the next call. Vetting an artifact before admission does not settle this. A safe variant and a leaking variant can produce the same admission evidence, and a sound gate then cannot relax that site for either. We make that condition precise, which leaves the last boundary a deployment can still act on. We present Provenance-Aware Capability Enforcement (PACE), which mediates every tool call immediately before it executes. Path confinement proposes an executable cut of represented influence paths, while capability and effect verification checks schema-defined effects against authority compiled from the authenticated request. We distinguish the certified execution contract from the evaluated configuration, which can restore an authorized call after a proposed block or apply a declared repair. Confinement requires the final action to preserve the certified cut. On eight executable agent-security benchmarks with three target-model families, the evaluated configuration gives strictly lowest attack success in 62 of 79 eligible attack columns and ties in 14; full-benchmark native utility loses at most three points relative to the undefended agent. A complete ablation over 1167 paired cases attributes most security gains to effect verification and refusal control to boundary adaptation. A reduced-scale adaptive search succeeds on 0/30 out-of-authority targets against the defense.
Figures & tables
Figure 1: One mediated step of PACE . Phase I–PROPOSE ( Section 4.1 ) freezes the call, extends the graph, and expands the effect atoms; phase II–CUT & CERTIFY ( Section 4.2 ) cuts and certifies against committed evidence alone; phase III–ENFORCE ( Section 4.3 ) installs the manifest and dispatches; phase IV–FINALIZE ( Section 4.3 ) commits the outcome.
Qwen
DeepSeek
ChatGPT
Method
ASR ↓
Util.
ASR ↓
Util.
ASR ↓
Util.
AgentDojo
Baseline
48.0
62.59
99.84
75.68
100.0
73.77
CaMel
0.32
35.14
28.14
60.04
36.09
57.02
DTA
0.95
29.84
33.33
54.0
55.01
56.01
DataSent.
35.29
73.77
78.06
85.0
84.1
82.95
Table 1: Main results on benchmarks with target-model families. ASR is the worst attack group for that benchmark and model. Each block shows the undefended agent, PACE , and the three strongest published defenses, ranked by mean worst-group ASR. Util. is the benchmark’s own score and is higher better except on MCPTox, where it is the refusal rate. Best ASR per block and model in bold.
A0
A1
A2
A3
A4
A5
A6
A7
Path confinement (P)
✓
✓
✓
✓
Capability and effect (C)
✓
✓
✓
✓
Boundary adaptation (B)
✓
✓
✓
✓
Attack success (%), lower better
WASP, plain text
8.3
8.3
0.0
0.0
8.3
0.0
8.3
0.0
WASP, URL injection
8.3
8.3
8.3
0.0
0.0
0.0
16.7
0.0
Table 2: Ablation study with the local Qwen-3.8-27B endpoint.
Appendix figures & tables14 assets
Supplementary material from the paper’s appendix.
Appendix
Notation
Description
Agent execution and threat model ( Section 2.1 )
M
Frozen policy that alternates free text with tool calls
T
Registered tool set
at
Tool call proposed at step t
ot
Observation returned by at
r
Answer returned to the user, itself an egress site
Appendix
Table 3: List of notations and their descriptions used in the paper.
Field
Value
Target models
gpt-5.6-luna ; DeepSeek-V4-Flash ; local Qwen-3.8-27B served by vLLM
Decoding
Temperature 0 ; maximum completion length 1024 tokens
Ablation seed
20260828 ; single deterministic paired run
Concurrency
8 workers; throughput only, no effect on the statistical population
Defense lanes
none , pace , and the external baselines of Table 7
Layer switches
Path confinement, capability and effect verification, and execution-boundary adaptation, toggled independently for the ablation arms of Table 6
Appendix
Table 4: Frozen evaluation and policy configuration.
Benchmark
Scaffold and native evaluator
Paired cases
AgentDojo
Official task, security, and utility functions
198
AgentDyn
Official three-domain runner and utility scorer
60
WASP
Official end-to-end evaluator; plain-text and URL groups
24
ASB
Official agents and metric scripts; seven run variants
280
PASB
Official personalized-agent workflow; IPI and four memory subsets
58
InjecAgent
Official tools and cases; Base and Enhanced splits
212
Appendix
Table 5: Benchmarks, native evaluators, and the paired-subset size used in the ablation. Commit and data revisions are recorded in the release manifest of Section E.6 .
Arm
P
C
B
Role
A0
Wrapper-matched no-defense corner
A1
✓
Path confinement alone
A2
✓
Capability and effect verification alone
A3
✓
Execution-boundary adaptation alone
A4
✓
✓
Path confinement with verification
A5
✓
✓
Path confinement with boundary adaptation
Appendix
Table 6: Ablation arms. P is path confinement, C is capability and effect verification, B is execution-boundary adaptation. All eight arms run on the same paired subset, so component contrasts are paired; zero-ASR floors can prevent identifying a security benefit.
Family
Methods
Benchmarks and fidelity
No defense
Baseline, ReAct
All eight; native
Trusted flow and IFC
CaMeL ( Debenedetti et al., 2025 ) , FIDES ( Costa et al., 2025 ) , DRIFT ( Li et al., 2025a )
AgentDojo, AgentDyn, WASP, PASB, MSB; native on AgentDojo, validated port elsewhere
Tool-use safeguards
Progent ( Shi et al., 2025 ) , Tool Filter ( Debenedetti et al., 2024 ) , Tool Allowlist ( Wang et al., 2026c ) , MCIP Guardian ( Jing et al., 2025 ) , ToolShield ( Li et al., 2026b )
AgentDojo, ASB, AgentDyn, MCPTox, MSB, PASB; native or validated port
Re-execution
MELON ( Zhu et al., 2025 )
AgentDojo; native
Detection
DataSentinel ( Liu et al., 2025 ) , PI-Detector ( Debenedetti et al., 2024 ) , PIGuard ( Li et al., 2025b ) , PromptGuard-2 ( Llama Team, 2025 ) , Metadata Sanitization ( Wang et al., 2026c )
AgentDojo, InjecAgent, ASB, AgentDyn, MCPTox, PASB, MSB; pre-filter plus end-to-end
Prompt-level
Delimiters, Sandwich, Spotlighting, Instructional Prevention, Direct and PoT Paraphrase, PoT Shuffle, Repeat User Prompt
WASP, ASB, InjecAgent, AgentDyn, PASB; native
Appendix
Table 7: External baselines by family, with the benchmarks each is reported on and the fidelity of the port.
Method
Qwen-3.8-27B
DeepSeek-V4-Flash
gpt-5.6-luna
Plain-text
URL Injection
Utility
Plain-text
URL Injection
Utility
Plain-text
URL Injection
Utility
ASR ↓
ASR ↓
Accuracy ↑
ASR ↓
ASR ↓
Accuracy ↑
ASR ↓
ASR ↓
Accuracy ↑
Baseline
19.05
14.29
78.57
9.52
7.14
79.76
4.76
4.76
86.90
Prompt Filter
9.52
14.29
82.14
7.14
4.76
75.00
11.90
9.52
83.33
Simple Static
20.37
15.32
72.64
16.67
14.29
64.29
19.04
14.29
84.52
FIDES
2.38
40.48
0.0
19.05
30.95
32.14
4.76
23.81
59.52
Appendix
Table 8: Results on WASP under the official end-to-end evaluator. Attack success is reported for the plain-text and URL-injection formats; utility is the native user-task accuracy. Lower is better for ASR, higher for utility; best ASR per block in bold.
Method
Qwen-3.8-27B
DeepSeek-V4-Flash
gpt-5.6-luna
Direct
Import Instructions
Tool Knowledge
Utility
Direct
Import Instructions
Tool Knowledge
Utility
Direct
Import Instructions
Tool Knowledge
Utility
ASR ↓
ASR ↓
ASR ↓
Accuracy ↑
ASR ↓
ASR ↓
ASR ↓
Accuracy ↑
ASR ↓
ASR ↓
ASR ↓
Accuracy ↑
Baseline
48.00
30.21
30.52
62.59
97.62
99.68
99.84
75.68
99.05
99.84
100
73.77
CaMeL
0.32
0.32
0.32
35.14
20.51
24.48
28.14
60.04
31.00
36.09
34.02
57.02
DTA
0.48
0.95
0.95
29.84
24.17
28.46
33.33
54.00
42.13
48.01
55.01
56.01
Spotlighting
98.40
86.30
80.00
62.47
100
92.05
90.14
75.42
100
100
100
70.01
Appendix
Table 9: Results on AgentDojo under the official security and utility functions, by attack type. Lower is better for ASR, higher for utility; best ASR per block in bold.
Method
Qwen-3.8-27B
DeepSeek-V4-Flash
gpt-5.6-luna
Direct Harm
DS-S1
DS-S2
Data-stealing
Direct Harm
DS-S1
DS-S2
Data-stealing
Direct Harm
DS-S1
DS-S2
Data-stealing
ASR(%) ↓
ASR(%) ↓
ASR(%) ↓
ASR(%) ↓
ASR(%) ↓
ASR(%) ↓
ASR(%) ↓
ASR(%) ↓
ASR(%) ↓
ASR(%) ↓
ASR(%) ↓
ASR(%) ↓
Baseline
12.35
23.71
79.07
18.75
19.61
28.13
52.29
14.71
12.55
18.57
50.50
9.38
ReAct
24.31
27.94
92.11
25.74
27.84
37.68
79.02
29.77
16.86
23.35
66.14
15.44
PromptGuard
13.14
24.45
80.45
19.67
21.57
31.80
64.74
20.59
10.00
17.28
56.38
9.74
Prompt-Sandwich
0.59
0.18
0.15
0.12
4.71
14.89
50.62
7.54
2.16
17.10
73.12
12.50
Appendix
Table 10: Results on the InjecAgent-Enhanced split. Direct Harm and the two data-stealing stages DS-S1 and DS-S2 are native attack-success rates, as is the aggregate data-stealing column. DS-S2 is conditional and excluded, together with the derived aggregate, from the main-table summary and 79-column comparison. Lower is better; best per block in bold.
Method
Qwen-3.8-27B
DeepSeek-V4-Flash
gpt-5.6-luna
DPI
OPI
MP
Mixed
PoT
Utility
DPI
OPI
MP
Mixed
PoT
Utility
DPI
OPI
MP
Mixed
PoT
Utility
ASR(%) ↓
ASR(%) ↓
ASR(%) ↓
ASR(%) ↓
ASR(%) ↓
Accuracy(%) ↑
ASR(%) ↓
ASR(%) ↓
ASR(%) ↓
ASR(%) ↓
ASR(%) ↓
Accuracy(%) ↑
ASR(%) ↓
ASR(%) ↓
ASR(%) ↓
ASR(%) ↓
ASR(%) ↓
Accuracy(%) ↑
Baseline
84.92
66.00
13.25
88.50
56.25
30.29
74.08
58.25
36.50
70.75
26.00
64.00
81.08
31.50
27.00
72.75
73.50
80.25
Delimiters
88.08
73.75
13.50
86.75
23.75
24.61
88.42
85.50
51.75
93.00
38.00
54.39
81.63
68.25
48.25
88.00
76.25
66.57
Sandwich
85.33
61.75
11.75
88.50
68.75
32.50
81.83
76.25
46.00
59.75
25.75
69.07
71.21
61.50
52.00
63.25
57.25
87.86
Instructional Prevention
71.92
60.75
17.00
82.00
52.75
33.93
58.75
61.75
40.75
78.50
20.75
72.50
54.57
40.75
43.25
83.75
46.00
87.25
Appendix
Table 11: Results on ASB across its five attack settings, with the native utility score. Lower is better for ASR, higher for utility; best ASR per block in bold.
Method
Qwen-3.8-27B
DeepSeek-V4-Flash
gpt-5.6-luna
Shopping
Github
DailyLife
Utility
Shopping
Github
DailyLife
Utility
Shopping
Github
DailyLife
Utility
ASR(%) ↓
ASR(%) ↓
ASR(%) ↓
Utility(%) ↑
ASR(%) ↓
ASR(%) ↓
ASR(%) ↓
Utility(%) ↑
ASR(%) ↓
ASR(%) ↓
ASR(%) ↓
Utility(%) ↑
Baseline
11.11
17.22
32.50
63.39
12.78
17.22
19.50
65.00
7.78
10.56
19.50
74.82
Repeat User Prompt
6.67
15.56
34.50
61.25
22.78
10.00
31.50
61.43
12.28
10.00
31.50
70.54
Spotlighting
3.33
13.89
5.50
65.00
13.33
7.22
3.00
58.21
17.22
15.00
27.50
64.46
Tool Filter
0.0
0.0
0.05
7.86
1.67
2.78
2.00
8.04
2.78
3.89
6.50
9.11
Appendix
Table 12: Results on AgentDyn across its three task domains, with the native utility score. Lower is better for ASR, higher for utility; best ASR per block in bold.
Method
Qwen-3.8-27B
DeepSeek-V4-Flash
gpt-5.6-luna
IPI
Short Read
Short Modify
Long Read
Long Modify
Utility
IPI
Short Read
Short Modify
Long Read
Long Modify
Utility
IPI
Short Read
Short Modify
Long Read
Long Modify
Utility
ASR(%) ↓
ASR(%) ↓
ASR(%) ↓
ASR(%) ↓
ASR(%) ↓
Accuracy(%) ↑
ASR(%) ↓
ASR(%) ↓
ASR(%) ↓
ASR(%) ↓
ASR(%) ↓
Accuracy(%) ↑
ASR(%) ↓
ASR(%) ↓
ASR(%) ↓
ASR(%) ↓
ASR(%) ↓
Accuracy(%) ↑
Baseline
0.0
85.00
62.50
62.50
75.00
98.47
0.0
67.50
72.50
60.00
92.50
99.24
2.29
55.00
42.50
50.00
97.5
100.0
Delimiters
0.0
25.00
0.0
45.00
20.00
99.23
0.0
30.00
55.00
42.50
32.50
98.47
1.53
22.50
35.00
30.00
35.00
99.24
Sandwich
0.0
85.00
70.00
65.00
72.50
99.24
0.0
72.50
85.00
57.50
87.50
98.47
3.82
40.00
25.00
42.50
82.50
100.0
Instructional Prevention
0.0
17.50
0.0
17.50
0.0
97.67
0.0
20.00
17.50
27.50
17.50
96.95
6.87
17.50
7.50
27.50
25.00
98.47
Appendix
Table 13: Results on PASB for indirect prompt injection and the four memory subsets, with the native IPI utility. Lower is better for ASR, higher for utility; best ASR per block in bold.
Method
Qwen-3.8-27B
DeepSeek-V4-Flash
gpt-5.6-luna
Template-1
Template-2
Template-3
Refusal
Template-1
Template-2
Template-3
Refusal
Template-1
Template-2
Template-3
Refusal
ASR(%) ↓
ASR(%) ↓
ASR(%) ↓
Refusal Rate(%) ↓
ASR(%) ↓
ASR(%) ↓
ASR(%) ↓
Refusal Rate(%) ↓
ASR(%) ↓
ASR(%) ↓
ASR(%) ↓
Refusal Rate(%) ↓
Baseline
22.39
8.84
32.52
0.16
56.54
52.61
56.06
0.42
23.04
26.10
86.93
27.96
Metadata Sanitization
5.94
1.38
17.41
0.15
13.09
9.91
28.60
0.37
5.76
5.43
91.48
26.67
PI-Detector
20.53
7.37
25.22
0.40
85.34
4.59
50.19
1.63
48.17
3.97
95.45
30.14
PI-Guard
11.59
3.09
11.07
0.91
50.79
27.56
30.68
3.15
34.03
9.19
63.64
46.32
Appendix
Table 14: Results on MCPTox for the three poisoning templates, with the native refusal rate. Lower is better for both; refusal is reported so that a low attack-success rate cannot be read without its cost. Best ASR per block in bold.
Method
Qwen-3.8-27B
DeepSeek-V4-Flash
gpt-5.6-luna
Plan
Call
Response
Multi-Stage
PUA
NRP
Plan
Call
Response
Multi-Stage
PUA
NRP
Plan
Call
Response
Multi-Stage
PUA
NRP
ASR(%) ↓
ASR(%) ↓
ASR(%) ↓
ASR(%) ↓
Accuracy(%) ↑
Accuracy(%) ↑
ASR(%) ↓
ASR(%) ↓
ASR(%) ↓
ASR(%) ↓
Accuracy(%) ↑
Accuracy(%) ↑
ASR(%) ↓
ASR(%) ↓
ASR(%) ↓
ASR(%) ↓
Accuracy(%) ↑
Accuracy(%) ↑
Baseline
5.67
38.75
26.35
18.70
46.34
36.29
61.33
100.0
29.68
52.00
91.24
49.29
67.00
98.75
9.03
50.20
89.08
54.51
MCIP Guardian
13.33
50.00
14.03
8.40
59.15
55.70
8.00
77.50
26.13
37.30
96.55
51.32
29.33
76.25
8.55
36.90
92.34
56.87
PI-Detector
16.67
38.75
7.02
5.20
54.36
49.83
55.33
53.75
18.55
48.40
94.81
51.18
65.00
70.00
7.10
47.10
90.85
56.22
PI-Guard
14.00
37.50
13.87
9.00
56.42
53.26
32.00
57.50
36.94
36.90
91.98
49.95
40.67
66.25
10.00
38.50
89.97
54.25
Appendix
Table 15: Results on MSB by attack stage, with the native PUA and NRP scores. Attack success is lower better; PUA and NRP are higher better. Cells marked -- were not produced by the official scorer for that configuration.
Quantity
Statistic
Interpretation
forward_simulation_ observed=true
Rate + CI
Last event required no graph expansion; not proof of H1
op.unknown expansions
Histogram/quantiles
Observed abstraction incompleteness
Unbindable minimum cuts
Rate + CI
Graph separation lacking an executor action
Stale-state rejection
Rate + CI
Freshness check exercised before dispatch
Effect-obligation failures
Histogram by obligation
Which of the nine checks fails, and how often
Bound action executed
Rate + CI
Observable executor agreement relevant to H4
Appendix
Table 16: Certificate-health quantities recorded per episode.
Tool-using LLM agents must act on untrusted webpages, emails, files, and API outputs while issuing privileged tool calls. Existing defenses often mediate trust at the granularity of an entire tool invocation, forcing a brittle choice in mixed-trust workflows: allow external content to influence a call and risk hijacked destinations or commands, or quarantine the call and block benign retrieval-then-act behavior. The key observation behind this paper is that indirect prompt injection becomes dangerous not when untrusted content appears in context, but when it determines an authority-bearing argument. We present \textsc{PACT} (\emph{Provenance-Aware Capability Contracts}), a runtime monitor that assigns semantic roles to tool arguments, tracks value provenance across replanning steps, and checks whether each argument's origin satisfies its role-specific trust contract. Under oracle provenance, \textsc{PACT} achieves 100% utility and 100% security on mixed-trust diagnostic suites, while flat invocation-level monitors incur false positives or false negatives. In full AgentDojo deployments across five models, \textsc{PACT} reaches 100% security on the three strongest models while recovering 38.1--46.4% utility, 8--16 percentage points above CaMeL at the same security level. Ablations show that both semantic roles and cross-step provenance are necessary. \textsc{PACT} reframes agent security as authority binding, and isolates the remaining deployment bottleneck to provenance inference and contract synthesis.
Linfeng Fan, Ziwei Li, Yuan Tian +3
Gaoling School of Artificial Intelligence, Renmin University of China, Beijing, China · King Abdullah University of Science and Technology, Thuwal, Saudi Arabia · Dongbei University of Finance and Economics, Dalian, China +1
Tool-using LLM agents remain vulnerable to indirect prompt injection because trusted instructions and untrusted observations share one context, allowing malicious content to steer consequential input-filtering defenses. Multi-path consensus defenses still leave a high attack success rate because they examine content or aggregated outputs rather than authorizing effects, especially for the within-tool attack, which preserves the intended tool but manipulates its arguments. Data-Flow Control such as CaMeL provides stronger guarantees, but incurs substantial time latency that limits practical deployment. We introduce ToolFence, which compiles a typed authorization blueprint before execution, enforces it through a deterministic monitor, and when the blueprint is incomplete asks a judge to grant new capabilities rather than adjudicate each concrete call. ToolFence provides two key advantages. First, its fine-grained provenance-aware authorization enables the system to distinguish user-authorized values from untrusted observations, effectively addressing the within-tool attack. Second, its deterministic fast path and capability-level runtime grants substantially reduce the frequency of expensive judge calls, improving runtime efficiency. On AgentDojo with Qwen3-max, ToolFence reduces overall ASR to near zero with only a 3.80 percentage-point clean-utility drop and practical runtime overhead.
Yanjie Li, Xiangyu He, Xuelong Dai +1
Hong Kong Polytechnic University · Shandong University
LLM agents increasingly rely on external tools, expanding capability while creating a new security boundary: third-party tools may appear benign at the interface level while embedding unsafe behavior in implementation. Existing defenses rely on weak metadata, collapse characterization and policy judgment into a single decision, or use heuristic/LLM enforcement that lacks deterministic, auditable reasoning over task context and multi-tool composition. This paper presents ToolGuardian, a policy-driven framework for securing agent-tool interactions through pre-admission vetting and task-aware runtime authorization. ToolGuardian uses progressive characterization to convert evidence into structured facts: descriptions capture declared intent, system-call traces expose coarse behavior, mock execution reveals observed effects, and source analysis identifies latent behavior. ToolGuardian's core contribution is an Answer Set Programming (ASP)-based declarative policy layer that reasons explicitly over capabilities, effects, task context, and composition. We compare ASP against heuristic and LLM-based policy realizations using identical inputs and output contracts. We evaluate ToolGuardian on 16 MCP-style tools, including 8 malicious variants derived from real open-source tools, and 20 runtime scenarios. For vetting, ASP reaches a deny-class F1 of 0.86 and 88% accuracy using description, syscall, and observed-effect evidence. For runtime authorization, fully specified realizations classify all scenarios correctly, while ablations show that removing compositional and conformance rules substantially degrades performance.
Arun Ravindran, Saurabh Deochake
Department of Electrical and Computer Engineering University of North Carolina at Charlotte Charlotte, NC · SentinelOne Mountain View, CA