Cheap, open agents make LLM pollution harder to mitigate
Authors: Raluca Rilla, Anne-Marie Nussberger, Rui Mata, Dirk U. Wulff
Organizations: Center for Humans and Machines, Max Planck Institute for Human Development, Berlin, Germany · International Max Planck Research School on Learning, Institutions, and Future Evolution (LIFE), Berlin, Germany · University of Basel, Basel, Switzerland · Center for Adaptive Rationality, Max Planck Institute for Human Development, Berlin, Germany · Vienna University of Economics and Business, Vienna, Austria
Large Language Model (LLM) pollution occurs when synthetic responses contaminate data intended to capture human behavior. High deployment costs have so far limited the risk posed by autonomous survey agents. However, open-weight models paired with open-source agentic frameworks may have removed this barrier. We compared the performance and detectability of nine agent configurations, ranging from fully open variants to closed commercial ones. Each agent autonomously completed a survey containing multiple response types yielding various detection checks. Fully open agents ran locally without usage fees and performed competitively with commercial alternatives. Open and commercial agents failed different sets of checks, and no single check reliably detected all agents, but open-text responses discriminated best between agents and humans. These findings identify fully open agents as a distinct risk for LLM pollution and support multilayered detection strategies emphasizing open-text analysis.
Figures & tables
Agent
Type
Time (median, s)
Failure rates (mean)
External tools
Behavior
Jailbreak
Prompt-leak
ttot
ta1
ta2
ta3
ta4
ta5
re- CAPTCHA
CF
Illusion
Box
AI-use
Copy
Paste
R-click
Drag
Drop
Beige
Off
Clip
Beige
Off
Clip
Devstral Small 2 24B
Open
626.50
3.64
3.60
3.27
3.29
3.33
0.08
0.00
0.08
0.60
0.00
0.00
0.00
0.00
0.00
0.00
0.00
0.00
1.00
0.00
0.00
0.00
Ministral 3 14B
305.00
3.04
3.24
3.14
3.18
2.97
0.30
0.00
0.08
0.50
0.00
0.00
0.00
0.00
0.00
0.00
0.00
0.00
1.00
0.00
0.00
0.00
Qwen 3 Coder 30B
554.50
3.61
3.18
3.18
3.24
3.08
0.10
0.00
0.43
0.70
0.00
0.00
0.00
0.00
0.00
0.00
0.00
0.00
1.00
0.00
0.00
0.00
Gemini 2.5 Flash
Mixed
145.00
3.16
3.40
3.35
3.37
3.26
0.00
0.00
0.23
0.95
0.40
0.00
0.00
0.00
0.00
0.00
0.00
0.00
1.00
0.00
0.00
0.00
Table 1: Agent performance across 40 runs per agent. Times are medians ( s ) for full survey ( ttot ) and for five association textboxes ( ta1 to ta5 ). Failure rates are the proportion of runs in which a check caught the agent. Cloudflare (CF) is an automated-traffic detection service. reCAPTCHA v3 scores each session from 0 (likely bot) to 1 (likely human); scores ≤ 0.5 fail. Illusion is an illusion-illusion item ( 10 ) . Box is an invisible checkbox asking respondents to confirm they have read instructions, failed when ticked. AI-use is failed by disclosing AI assistance. Copy, Paste, R-click, Drag, and Drop are failed by any attempt at these blocked actions. Jailbreak and prompt-leaking honeypots are hidden instructions to insert planted content or reveal the prompt, concealed as near-white text (Beige), outside the viewport (Off), or in a clipped element (Clip).
Figure 1: Standardized Euclidean distance of each agent configuration from the human response distribution; larger values indicate greater divergence. Semantic rows use embeddings of free-text answers: all five associations, the first and fifth associations individually, and the explanation of self-rated AI expertise. Behavioral rows cover total response time, time in the first association textbox, and counts of attempted paste, right-click, and copy actions. Self-report rows cover ratings of AI expertise, AI risk propensity, frequency of AI use, and trust in AI. Text items separate every configuration from humans by a wide margin, whereas behavioral and self-report items do not.
LLM-based agents are rapidly advancing, autonomously invoking external tools to complete multi-step tasks for users. However, agents often acquire more sensitive information than the task requires. Existing privacy benchmarks audit what the agent's response or outgoing actions disclose, but overlook the acquisition stage where data first enters the agent's context. The over-acquired information is then one careless action or one attack away from an outright leak. To assess its prevalence, we introduce \emph{PrivacyPeek}, a benchmark for evaluating acquisition-stage privacy leakage of LLM-based agents, with 1,182 cases across 7 acquisition behaviours and 16 application domains. Specifically, \emph{Acquisition Inspection} examines the agent's tool-call trajectory, both the tools it invokes and the data it receives, to detect when it acquires sensitive information beyond the task scope. \emph{Probe Elicitation} then issues a follow-up probe and measures how readily an attacker could elicit sensitive information the agent acquired but did not disclose. Our experiments on 10 LLM-based agents across 4 model families show that the unnecessary acquisition of sensitive information is widespread. In addition, we observe a correlation between the task-completion capability and acquisition-stage leakage. Prompt-level defences reduce only a small fraction of acquisition-stage leakage, leaving the majority unmitigated. These results make auditing acquisition-stage privacy both urgent and necessary. Our dataset and code are available at https://github.com/Xuan269/PrivacyPeek-Resource.
Mingxuan Zhang, Jiahui Han, Dadi Guo +5
Shanghai Artificial Intelligence Laboratory · Southeast University
LLM safety evaluations predominantly test models in isolation, yet deployed AI agents increasingly operate within persistent social environments alongside other agents. We introduce a Moltbook-style simulation platform where thousands of LLM agents interact across communities over a simulated month, and use it to evaluate privacy as a downstream safety concern under varying degrees of social pressure. We find that shifting from single turn to multi turn social evaluation amplifies privacy violations (CIMemories 19.95% to Ours 45.30% across OpenAI models), that leakage is socially contagious, with agents 8 times more likely to disclose sensitive information after observing a peer do so, and that explicit privacy instructions reduce but do not eliminate this effect, leaving leakage rates above 37.8% even with safeguards. Our findings suggest that static chat based safety benchmarks systematically underestimate risks in agentic deployment, and that social context alone is sufficient to elicit sensitive disclosures that single turn evaluations would never surface.
LLM-agent defenses are typically evaluated one session at a time. In deployment, however, attacks can be distributed across independent agents, teams, and runtimes, leaving each local guardrail with only a sparse fragment. We formalize cross-agent asynchronous campaign attribution: linking sessions from the same latent adversarial campaign without shared runtime state, test-time campaign labels, or attacker identity oracles. We introduce Asynchronous Attribution Fingerprint Vectors (A2FV), a lightweight proxy-side reference protocol for scoring pairwise campaign similarity from proxy-observable tool-use, timing, and prompt residue. We also construct SCD-v1, a controlled persona-matched benchmark with benign traffic, isolated attacks, multi-session campaigns, matched non-oracle evasion, and leakage audits. On SCD-v1, A2FV achieves 0.82 pairwise AUC for campaign linking, while score-only adaptations of per-session detectors and chunked LLM judges remain near chance under the same task. The strongest fixed signal is carried by structural and stylometric residue, while timing is retained as a diagnostic channel for richer proxy traces. Crossed-style controls show that the signal is partly style-sensitive but not reducible to style alone. Static and dimension-aware non-oracle stress tests further show that pairwise separability persists under controlled evasion. These results establish cross-agent campaign attribution as a distinct evaluation layer for securing LLM agents in the wild.