cs.AISep 25, 2026

Cheap, open agents make LLM pollution harder to mitigate

Authors: Raluca Rilla, Anne-Marie Nussberger, Rui Mata, Dirk U. Wulff

Organizations: Center for Humans and Machines, Max Planck Institute for Human Development, Berlin, Germany · International Max Planck Research School on Learning, Institutions, and Future Evolution (LIFE), Berlin, Germany · University of Basel, Basel, Switzerland · Center for Adaptive Rationality, Max Planck Institute for Human Development, Berlin, Germany · Vienna University of Economics and Business, Vienna, Austria

Abstract

Large Language Model (LLM) pollution occurs when synthetic responses contaminate data intended to capture human behavior. High deployment costs have so far limited the risk posed by autonomous survey agents. However, open-weight models paired with open-source agentic frameworks may have removed this barrier. We compared the performance and detectability of nine agent configurations, ranging from fully open variants to closed commercial ones. Each agent autonomously completed a survey containing multiple response types yielding various detection checks. Fully open agents ran locally without usage fees and performed competitively with commercial alternatives. Open and commercial agents failed different sets of checks, and no single check reliably detected all agents, but open-text responses discriminated best between agents and humans. These findings identify fully open agents as a distinct risk for LLM pollution and support multilayered detection strategies emphasizing open-text analysis.

Figures & tables

Explore similar work

CardsList
  1. PrivacyPeek: Auditing What LLM-Based Agents Acquire, Not Just What They Say

    May 29, 2026Mingxuan Zhang, Jiahui Han, Dadi Guo +5PrivacyAttacker Large Language Model

  2. Cross-Agent Campaign Attribution: Linking Asynchronous Attacks Across LLM Agents

    Jul 21, 2026SangJin Park, Myungsub Choi, Jineok Kim +1Multi-Llm AgentsLarge Language Model Agents