cs.AIOct 1, 2026

Beyond Final Accuracy: Auditing Communication in LLM Multi-Agent Systems

Authors: Shixuan Li, Wei Yang, Peiyu Zhang, Anzhe Cheng, Heng Ping, Paul Bogdan

Organizations: Ming Hsieh Department of Electrical and Computer Engineering University of Southern California, Los Angeles, CA 90089, USA · Thomas Lord Department of Computer Science University of Southern California, Los Angeles, CA 90089, USA

Abstract

Multi-agent communication aims to help agents benefit from one another's information. Yet improvements in system performance leave a fundamental ambiguity: do they reflect effective communication, a favorable agent architecture, or simply additional reasoning? Because communication methods are commonly evaluated within the systems they were designed for, these factors are difficult to disentangle. Final accuracy further merges corrected errors and corrupted answers into a single outcome, obscuring how communication changes decisions. We introduce Independent--Communicate--Revise (ICR), a controlled framework that evaluates communication as answer revision following independent reasoning. ICR fixes initial reasoning trajectories, measures correction and preservation conditional on both agents' initial correctness, and uses a no-message revision control to quantify gains beyond additional reasoning. Across four reasoning benchmarks, our audit of textual and latent communication reveals that similar aggregate accuracy can conceal substantially different revision behaviors. Compared with transmitting answers alone, full reasoning increases correction while reducing preservation on all four benchmarks, so richer messages amplify beneficial and harmful influence alike. Receiver-policy comparisons on MedQA and GPQA-D further show that a structured verification policy shifts every channel toward greater preservation and lower correction, while its effect on selectivity varies across channels and tasks. These findings challenge treating communication quality as an intrinsic property of a channel. ICR therefore recenters evaluation on selective revision, providing a unified framework for examining how message content and receiver policies jointly produce benefits and harms.

Figures & tables

Appendix figures & tables20 assets

Supplementary material from the paper’s appendix.

Appendix

Explore similar work

CardsList
  1. When Upstream Messages Override Correct Answers: A Controlled Study of Multi-Agent LLM Collaboration

    Sep 29, 2026Yaxin Gong, Gangyi Zhang, Chongming Gao +7Multi-Agent Large Language Model SystemsDownstream Reasoning

  2. Preventing Error Propagation in Multi-Agent AI through Runtime Monitoring

    Jun 27, 2026Shahnewaz Karim Sakib, Anindya Bijoy DasMulti-Agent ReasoningRuntime

  3. Do Latent Channels Actually Communicate? A Causal Audit of Latent Multi-Agent LLM

    Jul 29, 2026Huixiang Zhang, Mahzabeen EmuMulti-Agent Large Language Model SystemsModel Auditing