cs.LGSep 28, 2026

Emergence, Not Bandwidth: Physical Coupling and the Limits of Learned Multi-Agent Communication

Authors: Mihir Chauhan, Aniket Bera

Organizations: IDEAS Lab, Department of Computer Science Purdue University

Abstract

Rate-limited multi-agent teams raise three questions the emergent-communication literature has answered only empirically: what an optimal message should encode, what compression costs over a horizon, and when a learned protocol is unique enough for a teammate to read. We answer them for rate-limited Dec-POMDPs, then measure how far reinforcement learning falls short of the optimum. Our theorems fix what is achievable independently of any learner, so a gap between an engineered and a learned sender at the same bit budget is an optimization fact, not an information-theoretic one. We instantiate this on three MuJoCo arenas spanning zero, partial and rigid physical coupling, charging every condition exactly 2 bits per decision, and create the discriminating regime by closing a physical side channel within one arena, holding bodies, task and reward fixed. Communication value is governed by coupling: under rigid coupling through a shared object, no channel beats silence (+0.001 +/- 0.001, p = 0.982, n = 25), since proprioception already carries that information; without coupling, every condition solves the task; under partial coupling, the engineered 2-bit sender reaches an interquartile mean of 1.000 but the learned one reaches 0.482, indistinguishable from silence (p = 0.400, n = 25). With a shared alphabet, bandwidth cannot explain the gap. Warm-starting from an engineered receiver localizes the failure: the same channel reaches 0.857 versus 0.562 cold-started (p < 0.001), so it is neither representational nor one of maintenance; reinforcement learning fails to discover the protocol. Cross-play shows learned protocols are individually meaningful but mutually unintelligible: self-play 0.980 collapses to 0.144 across seeds, and our best constructed alignment leaves at least 77% of that gap. All headline results use 25 seeds per arena and seven published baselines at matched rate.

Figures & tables

Appendix figures & tables4 assets

Supplementary material from the paper’s appendix.

Appendix

Explore similar work

CardsList
  1. Decoupling Communication from Policy: Robust MARL under Bandwidth Constraints

    May 20, 2026Alexi Canesse, Benoît Goupil, Jesse Read +1Multi-Agent Reinforcement LearningBandwidth Extension

  2. MUTE: Return-Preserving Communication Unlearning for Efficient Multi-Agent Coordination

    Jul 3, 2026Rui Zuo, Qinwei Huang, Mingyang Li +3Multi-Agent Reinforcement LearningExact Unlearning

  3. Robust and Efficient Communication for Multi-Agent Learning

    Sep 14, 2026Rafael Pina, Varuna De Silva, Corentin ArtaudMulti-Agent Reinforcement LearningArtificial Intelligence Agents