Cheap, open agents make LLM pollution harder to mitigate
Authors: Raluca Rilla, Anne-Marie Nussberger, Rui Mata, Dirk U. Wulff
Organizations: Center for Humans and Machines, Max Planck Institute for Human Development, Berlin, Germany · International Max Planck Research School on Learning, Institutions, and Future Evolution (LIFE), Berlin, Germany · University of Basel, Basel, Switzerland · Center for Adaptive Rationality, Max Planck Institute for Human Development, Berlin, Germany · Vienna University of Economics and Business, Vienna, Austria
Large Language Model (LLM) pollution occurs when synthetic responses contaminate data intended to capture human behavior. High deployment costs have so far limited the risk posed by autonomous survey agents. However, open-weight models paired with open-source agentic frameworks may have removed this barrier. We compared the performance and detectability of nine agent configurations, ranging from fully open variants to closed commercial ones. Each agent autonomously completed a survey containing multiple response types yielding various detection checks. Fully open agents ran locally without usage fees and performed competitively with commercial alternatives. Open and commercial agents failed different sets of checks, and no single check reliably detected all agents, but open-text responses discriminated best between agents and humans. These findings identify fully open agents as a distinct risk for LLM pollution and support multilayered detection strategies emphasizing open-text analysis.
Figures & tables
Agent
Type
Time (median, s)
Failure rates (mean)
External tools
Behavior
Jailbreak
Prompt-leak
ttot
ta1
ta2
ta3
ta4
ta5
re- CAPTCHA
CF
Illusion
Box
AI-use
Copy
Paste
R-click
Drag
Drop
Beige
Off
Clip
Beige
Off
Clip
Devstral Small 2 24B
Open
626.50
3.64
3.60
3.27
3.29
3.33
0.08
0.00
0.08
0.60
0.00
0.00
0.00
0.00
0.00
0.00
0.00
0.00
1.00
0.00
0.00
0.00
Ministral 3 14B
305.00
3.04
3.24
3.14
3.18
2.97
0.30
0.00
0.08
0.50
0.00
0.00
0.00
0.00
0.00
0.00
0.00
0.00
1.00
0.00
0.00
0.00
Qwen 3 Coder 30B
554.50
3.61
3.18
3.18
3.24
3.08
0.10
0.00
0.43
0.70
0.00
0.00
0.00
0.00
0.00
0.00
0.00
0.00
1.00
0.00
0.00
0.00
Gemini 2.5 Flash
Mixed
145.00
3.16
3.40
3.35
3.37
3.26
0.00
0.00
0.23
0.95
0.40
0.00
0.00
0.00
0.00
0.00
0.00
0.00
1.00
0.00
0.00
0.00
Table 1: Agent performance across 40 runs per agent. Times are medians ( s ) for full survey ( ttot ) and for five association textboxes ( ta1 to ta5 ). Failure rates are the proportion of runs in which a check caught the agent. Cloudflare (CF) is an automated-traffic detection service. reCAPTCHA v3 scores each session from 0 (likely bot) to 1 (likely human); scores ≤ 0.5 fail. Illusion is an illusion-illusion item ( 10 ) . Box is an invisible checkbox asking respondents to confirm they have read instructions, failed when ticked. AI-use is failed by disclosing AI assistance. Copy, Paste, R-click, Drag, and Drop are failed by any attempt at these blocked actions. Jailbreak and prompt-leaking honeypots are hidden instructions to insert planted content or reveal the prompt, concealed as near-white text (Beige), outside the viewport (Off), or in a clipped element (Clip).
Figure 1: Standardized Euclidean distance of each agent configuration from the human response distribution; larger values indicate greater divergence. Semantic rows use embeddings of free-text answers: all five associations, the first and fifth associations individually, and the explanation of self-rated AI expertise. Behavioral rows cover total response time, time in the first association textbox, and counts of attempted paste, right-click, and copy actions. Self-report rows cover ratings of AI expertise, AI risk propensity, frequency of AI use, and trust in AI. Text items separate every configuration from humans by a wide margin, whereas behavioral and self-report items do not.