cs.AIOct 6, 2026

Where Rules End and Judges Begin: Measuring the Judgment Boundary in Multi-Agent Systems Security

Authors: Shaswata Mitra, Raj Patel, Subash Neupane, Sudip Mittal, Md Rayhanur Rahman, Shahram Rahimi

Organizations: Department of Computer Science The University of Alabama, Tuscaloosa, USA

Abstract

LLM-based multi-agent systems (MAS) engage tools, share memory, and delegate tasks, often encountering adversarial content. Current defenses for MAS are typically evaluated in isolation, focusing on one attack type at a time, which can lead to costly and hard-to-audit outcomes. This study organizes defenses into five principles, implementing them as DEFER1 (DEterministic-First Enforcement with Residual judgment), which includes a cascade of 28 checks that blocks what it can and refers the rest to a panel of four judges. In independent testing across four domains, attack success rates drop from about 30.0% to approximately 3.0%, with 78% of blocked attacks handled by deterministic checks. Only a quarter of proposals reach the judges in the security-operations domain, illustrating that the rules provide security for attacks violating clear policies, while judges manage those that only misrepresent intent. Both systems have weaknesses, such as a risk-score approval gate that inaccurately approves most attack proposals but few legitimate ones, highlighting the challenges in assessing threats accurately.

Figures & tables

Appendix figures & tables37 assets

Supplementary material from the paper’s appendix.

Appendix

Explore similar work

CardsList
  1. SoK: When Safe Agents Fail Together: The Security of Multi Agent LLM Systems

    Sep 1, 2026Rui Yang, Junjie Xu, Zhengyu Liu +4Model-Based Multi-Agent SystemsMulti-Agent Large Language Model Systems

  2. MAS-Shield: A Defense Framework for Secure and Efficient LLM MAS

    Nov 28, 2025Kaixiang Wang, Zhaojiacheng Zhou, Bunyod Suvonov +8LLM Defense MechanismsLinguistics

  3. The Capability Paradox: How Smarter Auditors Make Multi-Agent Systems Less Secure

    May 17, 2026Qiqi Liu, Thorsten Holz, Shilin Ye +1Attack-Success RateHijacking