cs.AIOct 5, 2026

AgentPrivArena: Evaluating and Auditing Real-world AI Agent Privacy

Authors: Shouju Wang, Haopeng Zhang

Organizations: University of North Carolina at Charlotte

Abstract

The rapid advancement of LLM agents has enabled systems to autonomously perform complex tasks through external tools, but their growing access to personal data introduces significant privacy risks. Existing benchmarks primarily evaluate LLM agent privacy through simulated trajectories and outcome-based metrics, limiting their ability to capture privacy risks arising during multi-step agent execution. In this work, we introduce AgentPrivArena, a framework for evaluating privacy risks in realistic LLM agent workflows. AgentPrivArena integrates authentic MCP tools and self-hosted services within a reproducible execution environment. We further propose trajectory-level privacy metrics that quantify unnecessary information access beyond final response leakage. Building on this framework, we introduce AgentPrivAudit, a runtime auditing approach for monitoring privacy violations during agent execution. Extensive experiments on state-of-the-art LLM agents reveal substantial privacy risks overlooked by existing evaluation paradigms, highlighting the importance of trajectory-level auditing for trustworthy agent deployment.

Figures & tables

Appendix figures & tables5 assets

Supplementary material from the paper’s appendix.

Appendix

Explore similar work

CardsList
  1. PrivacyPeek: Auditing What LLM-Based Agents Acquire, Not Just What They Say

    May 29, 2026Mingxuan Zhang, Jiahui Han, Dadi Guo +5PrivacyAttacker Large Language Model

  2. ToolPrivacyBench: Benchmarking Purpose-Bound Privacy in Tool-Using LLM Agents

    Jun 26, 2026Shijing Hu, Liang Liu, Zhu Meng +1Language-Model AgentsPrivacy