The rapid advancement of LLM agents has enabled systems to autonomously perform complex tasks through external tools, but their growing access to personal data introduces significant privacy risks. Existing benchmarks primarily evaluate LLM agent privacy through simulated trajectories and outcome-based metrics, limiting their ability to capture privacy risks arising during multi-step agent execution. In this work, we introduce AgentPrivArena, a framework for evaluating privacy risks in realistic LLM agent workflows. AgentPrivArena integrates authentic MCP tools and self-hosted services within a reproducible execution environment. We further propose trajectory-level privacy metrics that quantify unnecessary information access beyond final response leakage. Building on this framework, we introduce AgentPrivAudit, a runtime auditing approach for monitoring privacy violations during agent execution. Extensive experiments on state-of-the-art LLM agents reveal substantial privacy risks overlooked by existing evaluation paradigms, highlighting the importance of trajectory-level auditing for trustworthy agent deployment.
Figures & tables
Figure 1: AgentPrivArena. Left: each source scenario is converted into a task specification plus per-service seed files, so the records it describes exist as service state rather than as text in a prompt. Centre: the seeded services expose their native APIs through MCP servers, and the sensitive records sit inside them, reachable only by the agent’s own reads. Right: the agent plans, calls tools and accumulates observations until it commits a final action. The highlighted observation is the reason a single-step evaluation is not enough: task-relevant content (green) and protected content (red) arrive together in one read and persist in context from then on.
Work
CI
Exec.
Live env.
Trace
PrivacyLens ( 2024 )
✓
✗
✗
✗
AgentDAM ( 2025 )
–
✓
–
✗
AgentLeak ( 2026 )
✓
✗
✗
✓
CI-Work ( 2026 )
✓
✗
✗
✗
MPCI-Bench ( 2026 )
✓
✗
✗
✗
PiSAs ( 2026 )
✓
✗
✗
✗
Table 1: Agentic contextual privacy evaluations, CI : grounded in CI theory. Exec. : the agent runs with real tool calls. Live env. : provide realistic agent environment. Trace : trajectory level evaluation.
Application function
Service
Version
Knowledge base
BookStack
26.03.3
Direct messaging
Mattermost
11.6.0
Team chat
Rocket.Chat
6.13
Email
Mailpit
1.29.7
Social media
GoToSocial
0.21.2
Calendar scheduling
Radicale
3.6.1
Table 2: The local services AgentPrivArena runs to instantiate common privacy-sensitive application functions as executable agent tasks. All six are open-source projects, self-hosted in containers, and Appendix B gives each project’s home page and image digest prefix.
Executor
Complete ↑
Clarify
Err. rec. ↑
GPT-5.4
100.0
1.5
100.0
DeepSeek-V4-Pro
99.2
1.3
100.0
Kimi-K2.6
99.5
3.1
100.0
Mistral-Large-3
97.9
4.4
93.3
Gemini-Flash-2.5
98.2
4.6
99.1
Table 3: Execution fidelity on the 389 executable tasks.
Figure 2: AgentPrivAudit attaches at two boundaries, and the vertical spines separate what executes from what audits. Dashed edges carry audit information; solid edges are the execution path.
Leak rate (%, ↓ )
Helpfulness ( 0 – 3 , ↑ )
Executor
C0
C1
C2
C3
C4
C0
C1
C2
C3
C4
GPT-5.4
33.4
27.1
23.7
20.5
17.3
2.68
2.67
2.66
2.65
2.64
DeepSeek-V4-Pro
58.8
53.9
31.4
22.4
19.3
2.58
2.67
2.57
2.58
2.64
Kimi-K2.6
54.8
52.1
28.8
18.5
18.1
2.61
2.66
2.54
2.63
2.61
Mistral-Large-3
46.5
44.5
25.1
18.6
16.2
2.54
2.61
2.55
2.55
2.59
Gemini-Flash-2.5
40.3
37.6
23.5
18.7
18.1
2.56
2.59
2.43
2.51
2.48
Table 4: Outcome-level results. Both metrics score the same committed action, n=389 tasks per cell. Conditions: C0 no mitigation, C1 privacy prompt, C2–C4 the audit under PII, data minimization and contextual integrity. Avg is pooled over tasks. C4 is the lowest leak rate in every row; the underlined cell is the lowest helpfulness in the table.
Figure 3: The three trajectory-level privacy metrics, audited conditions. (a) Exposure rate , over all reference protected items. (b) Extraction recall , over the exposed ones: what the read boundary represented as a flow given that the agent had seen it. (c) Disposition profile , over extracted flows; only segments above 8% carry a label, so PII’s 1.2% Block share is unlabelled.
Figure 4: Leak rate per executor, by condition (codes as in Table 4 ). Each rule spans the five executors and the right-hand column gives its width in percentage points. Prompting leaves the spread where it found it; enforcement collapses it.
Figure 5: The same three criteria, stated as an instruction and enforced by the mechanism.
Figure 6: Two executors, four arms, criterion and boundaries fixed. Along a line is the auditor’s effect, between the lines the executor’s. Paired, the auditor swap gives −6.2 and −10.1 pp ( p<10−2 ), the executor swap ±1.8 pp ( p>0.3 ). Each arm uses 389 tasks; paired executor contrasts use 382 – 385 matched tasks.
Appendix figures & tables5 assets
Supplementary material from the paper’s appendix.
Appendix
Service
Discovery
Access
Write
BookStack
search_pages , list_pages
read_page
create_page , update_page , delete_page
Mattermost
search_messages † , list_users
read_messages
send_message
Rocket.Chat
search_messages † , list_channels , search_users
read_channel_history , get_user_info
send_channel_message
Mailpit
search_emails , list_contacts
read_email
send_email
GoToSocial
search_users , search_posts †
read_profile , read_user_posts
create_post
Radicale
search_events , list_events
read_event
–
Appendix
Table 5: The 28 tools the agent can call, by service and operation class. The write set is the one the platform enforces on: it is declared per service in the tool registry and is what the audit’s write boundary keys on. The discovery/access split describes intent rather than an information boundary, † marks discovery tools that return record content along with the identifier (Mattermost’s search returns whole message bodies, BookStack’s a snippet), which is why the audit treats every observation as extractable rather than trusting the tool class.
Service
Project and image
BookStack
https://www.bookstackapp.com
lscr.io/linuxserver/bookstack
sha256:3014bc25b8ce
Mattermost
https://mattermost.com
mattermost/mattermost-team-edition
sha256:88565cb169f3
Appendix
Table 6: Provenance of the local services. Digests are truncated to twelve hexadecimal characters.
Role
Name in the paper
Identifier
Executor
GPT-5.4
openai/gpt-5.4
DeepSeek-V4-Pro
openai/DeepSeek-V4-Pro
Kimi-K2.6
openai/Kimi-K2.6
Mistral-Large-3
openai/Mistral-Large-3
Gemini-Flash-2.5
gemini/gemini-2.5-flash
Auditor
GPT-5.4
openai/gpt-5.4
Appendix
Table 7: The three model roles. grok-4-20-non-reasoning appears only as an executor and the alternative auditor in § 6.2 ; it was chosen because it is the non-reasoning member of the pool, which is the axis that ablation varies. Decoding is left at each endpoint’s default in every role, so no temperature is tuned per condition. The exec × auditor cells of § 6.2 predate run_metadata , so their executor and auditor are recovered from the per-call model identifiers in the event log rather than from the directory name; all eight recovered assignments match.
Executor
Auditor
vs C0
in
out
in
out
C0
35.5k
1.5k
—
—
1.00×
C1
22.4k
1.5k
—
—
0.65×
C2
59.0k
1.8k
7.3k
10.6k
2.13×
C3
60.2k
1.7k
7.5k
12.2k
2.21×
C4
56.8k
1.7k
8.1k
10.0k
2.07×
Appendix
Table 8: Token cost per task. Auditing roughly doubles the total, but the auditor accounts for only 23 – 24% of it: most of the increase is the executor working over a longer context once the inventory and steering note are in it. The privacy prompt is cheaper than no mitigation, because the agent it produces explores less.
Setting
n
Read steps
Deep
mean
med.
max
( ≥8 )
Static PrivacyLens
493
1.9
2
8
0.2%
Live, GPT-5.4
389
6.8
6
32
30.6%
Live, DeepSeek-V4-Pro
389
5.9
6
16
20.3%
Live, Kimi-K2.6
389
5.2
5
22
15.4%
Live, Mistral-Large-3
389
3.9
4
11
3.6%
Appendix
Table 9: Trajectory shape, static versus live. A step is one content-bearing read; static counts are Action: entries in the pre-authored trajectory. Deep tasks ( ≥8 reads) are the regime where contextual judgment becomes load-bearing: 14.3% of live executions against 0.2% of the static benchmark.