Communication Gain and Delay Cost Under Cross-Timestep Delays in Cooperative Multi-Agent Reinforcement Learning
Organizations: The State Key Laboratory for Manufacturing Systems Engineering School of Automation Science and Engineering, Xi’an Jiaotong University
Abstract
Communication is essential for coordination in \emph{cooperative} multi-agent reinforcement learning under partial observability, yet \emph{cross-timestep} delays cause messages to arrive multiple timesteps after generation, inducing temporal misalignment and making information stale when consumed. We formalize this setting as a delayed-communication partially observable Markov game (DeComm-POMG) and decompose a message's effect into \emph{communication gain} and \emph{delay cost}, yielding the Communication Gain and Delay Cost (CGDC) metric. We further establish a value-loss bound showing that the degradation induced by delayed messages is upper-bounded by a discounted accumulation of an information gap between the action distributions induced by timely versus delayed messages. Guided by CGDC, we propose \textbf{CDCMA}, an actor--critic framework that requests messages only when predicted CGDC is positive, predicts future observations to reduce misalignment at consumption, and fuses delayed messages via CGDC-guided attention. Experiments on no-teammate-vision variants of Cooperative Navigation and Predator Prey, and on SMAC maps across multiple delay levels show consistent improvements in performance, robustness, and generalization, with ablations validating each component.
Figures & tables
| Task | Difficulty | CDCMA | CoDe | DACOM | TGCNet | T2MAC | SMS | G2ANet | ATOC |
|---|---|---|---|---|---|---|---|---|---|
| Cooperative Navigation | easy | -1.70 (0.07) | -3.11(0.14) | -3.16(0.43) | -2.08(0.32) | -2.23(0.04) | -2.57(0.11) | -2.24(0.09) | -2.51(0.62) |
| medium | -1.80 (0.36) | -3.09(0.05) | -3.14(0.13) | -2.21(0.35) | -2.27(0.08) | -2.67(1.01) | -2.48(0.07) | -2.43(0.14) | |
| hard | -1.78 (0.22) | -4.36(2.55) | -3.52(1.02) | -2.35(0.34) | -2.51(0.44) | -2.81(0.47) | -2.71(0.07) | -2.75(0.22) | |
| super_hard | -1.89 (0.13) | -4.42(2.75) | -3.00(0.21) | -2.56(0.36) | -2.53(0.17) | -3.04(0.64) | -2.94(0.06) | -2.79(0.45) | |
| Predator Prey | easy | -0.93 (0.03) | -1.82(0.03) | -1.97(0.23) | -1.53(0.14) | -1.10(0.15) | -1.46(0.27) | -1.17(0.01) | -2.38(0.38) |
| medium | -0.91 (0.04) | -1.83(0.03) | -2.19(0.42) | -1.60(0.14) | -1.27(0.22) | -1.36(0.17) | -1.40(0.05) | -2.75(0.54) |
| Method | Train | Test | ||
|---|---|---|---|---|
| easy | medium | hard | super_hard | |
| CDCMA | ||||
| CoDe | ||||
| DACOM | ||||
| TGCNet | ||||
| T2MAC | ||||
Appendix figures & tables9 assets
Supplementary material from the paper’s appendix.
Appendix
| Category | Symbol | Meaning |
| Delayed communication semantics | Joint environment–communication action | |
| Message sent from agent to agent at timestep | ||
| Message from agent that is available to agent at receiver timestep | ||
| Timely reference condition for sender at receiver timestep | ||
| Null / masked message input | ||
| Sampled communication delay for link at timestep |
| Hyperparameter | MPE | SMAC |
| Discount factor | 0.96 | 0.99 |
| Actor learning rate | ||
| Critic learning rate | ||
| Actor hidden dimension | 64 | 64 |
| Critic hidden dimension | 128 | 128 |
| Message dimension | 64 | 64 |
| Task | Low-gain | High-gain | ( ) | ( ) | Sign agreement (%) |
|---|---|---|---|---|---|
| PP | -0.019 | 0.088 | -0.013 | 0.076 | 79.1 |
| 1o_10b_vs_1r | -0.012 | 0.121 | -0.018 | 0.109 | 82.1 |