Rate-limited multi-agent teams raise three questions the emergent-communication literature has answered only empirically: what an optimal message should encode, what compression costs over a horizon, and when a learned protocol is unique enough for a teammate to read. We answer them for rate-limited Dec-POMDPs, then measure how far reinforcement learning falls short of the optimum. Our theorems fix what is achievable independently of any learner, so a gap between an engineered and a learned sender at the same bit budget is an optimization fact, not an information-theoretic one. We instantiate this on three MuJoCo arenas spanning zero, partial and rigid physical coupling, charging every condition exactly 2 bits per decision, and create the discriminating regime by closing a physical side channel within one arena, holding bodies, task and reward fixed. Communication value is governed by coupling: under rigid coupling through a shared object, no channel beats silence (+0.001 +/- 0.001, p = 0.982, n = 25), since proprioception already carries that information; without coupling, every condition solves the task; under partial coupling, the engineered 2-bit sender reaches an interquartile mean of 1.000 but the learned one reaches 0.482, indistinguishable from silence (p = 0.400, n = 25). With a shared alphabet, bandwidth cannot explain the gap. Warm-starting from an engineered receiver localizes the failure: the same channel reaches 0.857 versus 0.562 cold-started (p < 0.001), so it is neither representational nor one of maintenance; reinforcement learning fails to discover the protocol. Cross-play shows learned protocols are individually meaningful but mutually unintelligible: self-play 0.980 collapses to 0.144 across seeds, and our best constructed alignment leaves at least 77% of that gap. All headline results use 25 seeds per arena and seven published baselines at matched rate.
Figures & tables
condition
bits/decision
engineered (scripted sector symbol)
2.000
vq_k4 (learned, ours)
2.000
cat_k4 (Gumbel straight-through)
2.000
dru_b2 / stdru_b2
2.000
vq_k2 / vq_k8 / vq_k16
1.000 / 3.000 / 4.000
silence
0.000
Table 1: Communication rate on every arena, derived from each channel’s configuration and never from realized message statistics. Every finite charge is a fixed-length cost over the alphabet the channel actually uses. The engineered sender and the learned vq_k4 channel are exactly matched, as are three published baselines.
Figure 1: Communication value is set by physical coupling. Left: each condition’s success against silence, n=25 seeds per condition, 95%t -intervals. Right: the contrast of engineered minus learned, with both senders charged exactly 2 bits per decision. The matched-rate gap is negligible where the agents are rigidly coupled (proprioception already carries the message), small where they are uncoupled (any channel suffices), and large in between. Since the alphabets are identical by construction, the middle bar is not a bandwidth effect.
arena
condition
n
success
IQM [95% CI]
vs silence
p
uncoupled
oracle
25
1.000±0.000
1.000 [1.000, 1.000]
+1.000±0.000
—
engineered
25
0.999±0.001
1.000 [1.000, 1.000]
+0.999±0.001
<0.001
vq_k4
25
0.983±0.010
0.992 [0.977, 0.999]
+0.983±0.010
<0.001
silence
25
0.000±0.000
0.000 [0.000, 0.000]
reference
partial
oracle
25
0.984±0.032
1.000 [1.000, 1.000]
+0.490±0.124
<0.001
engineered
25
0.936±0.076
1.000 [1.000, 1.000]
+0.441±0.129
<0.001
Table 2: Success against silence on three arenas, n seeds per condition. IQM is the interquartile mean with a seed-level bootstrap interval.
method
source
n
success
vs ours ( p )
DIAL + VQ (ours)
Foerster et al. (2016)
25
0.983±0.010
—
RIAL
Foerster et al. (2016)
25
0.998±0.003
+0.016±0.011 (0.024)
ST-DRU
Vanneste et al. (2022)
25
0.998±0.002
+0.016±0.009 (0.010)
DRU
Foerster et al. (2016)
25
0.991±0.005
+0.009±0.011 (0.105)
Eccles biases
Eccles et al. (2019)
25
0.931±0.041
−0.052±0.040 (0.041)
AE-grounding
Lin et al. (2021)
25
0.931±0.043
−0.052±0.045 (0.050)
Table 3: Communication learners on the uncoupled arena. Same trunk, arena, rate and seeds; paired on seeds, Holm-corrected across the table. This is a positive control, not a ranking: all seven clear silence decisively and the whole field sits within 0.07 at the top of a saturated scale, which establishes that a rate-matched discrete channel of any of these kinds succeeds where coupling is zero. The two that beat our channel here are carried forward below to the arena that discriminates.
Appendix figures & tables4 assets
Supplementary material from the paper’s appendix.
Appendix
candidate meaning
NMI
lift
95% CI
p (Holm)
sector
0.232
+0.176
[+0.141,+0.213]
<0.001
bearing (4 buckets)
0.298
+0.167
[+0.137,+0.197]
<0.001
turn sign
0.070
+0.058
[+0.040,+0.077]
<0.001
range (4 buckets)
0.029
+0.021
[+0.011,+0.033]
0.002
elapsed time (4 buckets)
0.035
+0.011
[+0.006,+0.017]
0.002
aligned
0.218
+0.003
[+0.000,+0.010]
0.322
Appendix
Table 4: The code carries the goal sector and the required steering bearing in roughly equal measure. aligned — whether the vehicle is already pointed correctly — is the one candidate it does not track.
Figure 2: Per-seed evaluation success on the partially coupled arena, one point per seed, with the mean (bar) and interquartile mean (triangle) marked. The right panel histograms the learned channel alone. The paper reports an interquartile mean beside the mean on this arena because the distribution is bimodal; this is the evidence for that, and it also shows why the learned channel and silence are hard to separate: they have the same two-cluster shape, and the learned channel merely moves a few more seeds into the upper cluster. Values are those of Table 2 .
Figure 3: Cross-play success for every ordered pair of seeds on the uncoupled arena, one panel per alignment map. The bright diagonal in the leftmost panel is self-play; the dark field is every other pairing. Diagonal cells are gray in the alignment panels because a map from a seed to itself is not a pairing that was run. Panel labels give the cross-play mean and the share of the self-play gap the map recovers, both read from the same data as Section 4.5 .
Figure 4: Warm- against cold-started runs on the partially coupled arena, joined per seed, with the engineered and silent references drawn as horizontal lines. The steep bundle is the discovery effect: seeds that sit near the silent floor when cold-started reach the engineered level when handed a competent receiver. Seeds that do not improve are visible rather than absorbed into a mean. Values are those of Section 4.3 .
Communication enables coordination in multi-agent reinforcement learning (MARL), but many real-world applications, e.g., search-and-rescue with drone swarms, operate under severe bandwidth constraints. Many communication architectures still expose a coupled bottleneck in which a shared latent representation is used for both policy execution and inter-agent communication. Consequently, reducing message size directly limits the policy's latent space, often leading to significant performance degradation. We address this with two contributions. First, we introduce β, a normalised per-agent bandwidth budget that unifies sparsity, rounds, and message dimension into a single comparable constraint. Second, we provide SLIM, a minimal architecture that decouples the communication pathway from the policy's latent representation, allowing us to isolate the effect of bandwidth from the effect of policy capacity while benefiting from in-step communication. We evaluate our method on several partially-observable MARL benchmarks, where communication is essential. Our approach achieves state-of-the-art performance and exhibits scalability and robustness under limited communication, with only marginal degradation as bandwidth is reduced.
Alexi Canesse, Benoît Goupil, Jesse Read +1
École polytechnique (LIX), CNRS, Institut Polytechnique de Paris, Palaiseau, France
Inter-agent communication is critical for coordinating Multi-Agent Reinforcement Learning (MARL) agents under partial observability to perform effectively in cooperative games; however, real-world bandwidth constraints demand sparse interactions. Prior approaches primarily address this trade-off by optimizing information-theoretic surrogates. We argue that these statistical proxies are fundamentally misaligned with the true objective: a message can be highly informative yet irrelevant to the joint return of the task. In this work, we propose Message Unlearning for Targeted Efficiency (MUTE), a framework that views communication reduction as a value-guided machine unlearning problem. MUTE rigorously quantifies the Counterfactual Message Value using an attention-based estimator, and systematically unlearns the transmission of low-value messages from a policy trained without any communication constraints. This is achieved through a dual-objective mechanism that enforces communication sparsity while preserving the return of the original joint policy. We derive a theoretical upper bound on the performance gap induced by this sparsification, guaranteeing controlled return degradation. We also empirically evaluate MUTE on various complex multi-agent environments, achieving 80% to 90% bandwidth reduction while maintaining performance comparable to state-of-the-art baselines.
Rui Zuo, Qinwei Huang, Mingyang Li +3
Syracuse University · Air Force Research Laboratory
Effective communication is a cornerstone of distributed intelligence in Multi-Agent Reinforcement Learning (MARL), yet ensuring that generated messages are both informative and robust to physical constraints remains a significant challenge. This paper introduces Multi-Agent Regularized Communication (MARC), a novel framework inspired by information-theoretic principles of conditional mutual information. MARC employs an attention-based architecture coupled with a unique message regularization mechanism designed to minimize uncertainty regarding future system states, thereby inducing the learning of highly representative communication protocols. Crucially, we evaluate MARC under stringent communication bottlenecks and lossy channels, simulating the real-world constraints of autonomous robotic networks and decentralized systems. Our results demonstrate that MARC significantly outperforms state-of-the-art methods in complex cooperative domains. Furthermore, we provide a deep analysis of message characteristics, proving that MARC maintains high operational performance even under significant data compression, offering a scalable path for deploying intelligent agents in resource-constrained environments.
Rafael Pina, Varuna De Silva, Corentin Artaud
Institute for Digital Technologies · Loughborough University London, United Kingdom