cs.AIJul 31, 2026

More Debate, Same Evidence: Structural Limits of Homogeneous Multi-Agent Groundedness

Authors: Yuelyu Ji

Abstract

Large language model (LLM) judges are increasingly organized as multi-agent panels under the assumption that exchanging critiques improves judgment quality. We test this assumption for \emph{groundedness verification}, where a judge must determine whether a claim is supported by the supplied evidence. We evaluate a homogeneous three-agent panel on six public fact-verification and hallucination-detection benchmarks. Relative to a fixed single-agent reference, the panel's system-level accuracy difference ranges from +8.5+8.5 to 4.4-4.4 percentage points: two datasets show reliable gains, one shows a reliable loss, and three are statistically inconclusive. Because the reference and panel use different model variants, these differences characterize the complete systems rather than isolate a causal debate effect.

Explore similar work

CardsList