Learning-based multi-agent communication under limited bandwidth does not only require deciding what to communicate, but also structuring messages so that partial transmissions remain useful. We study this problem under prefix truncation, where only the first part of each message is received. To address it, we propose \textbf{AH-VIB}, an attention-based autoregressive variational communication model that combines a variational information bottleneck (VIB) with sequential message generation and a hierarchical robustness loss. We evaluate AH-VIB on a custom cooperative object-inspection and occupancy-mapping task, where agents equipped with a limited field-of-view sensor coordinate to scan inspection objects in an occupancy-grid world, under variable and fixed bandwidth conditions, and compare it against MADDPG, CommNet, a flat VIB baseline, and an autoregressive MLP ablation. AH-VIB achieves competitive mean return while improving performance reliability under the most constrained bandwidth conditions. These results indicate that AH-VIB improves the reliability and graceful degradation of learned communication under bandwidth constraints.
Figures & tables
Fig. 1: Cooperative object inspection under partial observability. Agents with limited fields of view progressively reveal free space and inspection-object surfaces while coordinating their coverage.
Hyperparameter
Value
Hyperparameter
Value
Hyperparameter
Value
Environment & Training
Optimisation
Exploration Noise
Total env. steps
5×105
Actor LR
3×10−4
Initial noise σ0
0.20
Episode length T
200
Critic LR
3×10−5
Decay factor
0.80
Replay buffer size
2×105
Discount γ
0.95
Decay over (steps)
1×105
Batch size
256
Target network τ
0.005
Min. noise σmin
0.05
Random exploration steps
4,000
Actor hidden layers
[512,256,256,128]
TABLE I: Training hyperparameters used in the experiments, shared across all algorithms unless noted.
Algorithm
VIB
AR
Hier.
Attn.
MADDPG
×
×
×
×
CommNet
×
×
×
×
VIB
✓
×
×
×
HVIB
✓
✓
✓
×
AH-VIB
✓
✓
✓
✓
TABLE II: Baseline design comparison.
ρ =0.10
ρ =0.25
ρ =0.75
ρ =1.0
AH-VIB
-0.2 ± 0.5
-0.2 ± 0.5
-0.4 ± 1.9
-0.4 ± 2.2
VIB
-7.7 ± 24.3
-6.2 ± 21.6
-6.1 ± 21.6
-7.3 ± 23.0
MADDPG
-5.4 ± 16.7
-5.3 ± 16.9
-13.7 ± 29.2
-1.1 ± 8.4
COMMNET
-1.8 ± 10.1
-2.1 ± 12.2
-0.1 ± 0.3
-0.1 ± 0.3
TABLE III: Episode return at selected bandwidth fractions (mean ± standard deviation). Bold indicates the best mean and lowest standard deviation in each column.
Communication enables coordination in multi-agent reinforcement learning (MARL), but many real-world applications, e.g., search-and-rescue with drone swarms, operate under severe bandwidth constraints. Many communication architectures still expose a coupled bottleneck in which a shared latent representation is used for both policy execution and inter-agent communication. Consequently, reducing message size directly limits the policy's latent space, often leading to significant performance degradation. We address this with two contributions. First, we introduce β, a normalised per-agent bandwidth budget that unifies sparsity, rounds, and message dimension into a single comparable constraint. Second, we provide SLIM, a minimal architecture that decouples the communication pathway from the policy's latent representation, allowing us to isolate the effect of bandwidth from the effect of policy capacity while benefiting from in-step communication. We evaluate our method on several partially-observable MARL benchmarks, where communication is essential. Our approach achieves state-of-the-art performance and exhibits scalability and robustness under limited communication, with only marginal degradation as bandwidth is reduced.
Alexi Canesse, Benoît Goupil, Jesse Read +1
École polytechnique (LIX), CNRS, Institut Polytechnique de Paris, Palaiseau, France
Cooperative multi-agent reinforcement learning (MARL) often relies on communication to mitigate partial observability, yet most existing protocols treat messages as flat dense vectors detached from the structure of the observations they summarize. This design overlooks an important source of inductive bias in many cooperative environments, where observations naturally follow a hierarchy such as groups and entities. We propose \textsc{HiComm}, a plug-in communication module that grounds messages in the sender's hierarchical observation. \textsc{HiComm} is receiver-driven: the receiver issues a query, and the hierarchy is resolved through a three-stage decoding process that first selects a group, then a sender, and then an entity within that group, returning the corresponding feature slice as the message. This converts communication from unstructured vector transmission into structured information retrieval over the sender's observation hierarchy. We instantiate this mechanism with Straight-Through Gumbel-Softmax for differentiable discrete selection and a lightweight shared projection design that attaches to standard MARL pipelines. Experiments across cooperative MARL tasks with different observation structures and coordination demands show that \textsc{HiComm} matches or outperforms representative learned communication baselines while reducing communication volume by up to 23× per receiver per episode.
Runze Zhao, Dongruo Zhou, Sumit Kumar Jha +2
Luddy School of Informatics, Computing, and Engineering Indiana University Bloomington Bloomington, IN 47408 · Department of Computer & Information Science & Engineering University of Florida Gainesville, FL 32611 · Department of Electrical Engineering and Computer Science United States Military Academy West Point, NY 10996 +1
Effective communication is a cornerstone of distributed intelligence in Multi-Agent Reinforcement Learning (MARL), yet ensuring that generated messages are both informative and robust to physical constraints remains a significant challenge. This paper introduces Multi-Agent Regularized Communication (MARC), a novel framework inspired by information-theoretic principles of conditional mutual information. MARC employs an attention-based architecture coupled with a unique message regularization mechanism designed to minimize uncertainty regarding future system states, thereby inducing the learning of highly representative communication protocols. Crucially, we evaluate MARC under stringent communication bottlenecks and lossy channels, simulating the real-world constraints of autonomous robotic networks and decentralized systems. Our results demonstrate that MARC significantly outperforms state-of-the-art methods in complex cooperative domains. Furthermore, we provide a deep analysis of message characteristics, proving that MARC maintains high operational performance even under significant data compression, offering a scalable path for deploying intelligent agents in resource-constrained environments.
Rafael Pina, Varuna De Silva, Corentin Artaud
Institute for Digital Technologies · Loughborough University London, United Kingdom