Hierarchical Multi-agent Reinforcement Learning for Warehouse Robot Coordination under Communication Loss
Authors: Weihao Sun, Gehui Xu, Andreas A. Malikopoulos
Organizations: Systems Engineering Program, Cornell University, Ithaca, NY 14850 USA · Department of Electrical and Electronic Engineering, Imperial College London, London SW7 2AZ, UK · Applied Mathematics, Systems Engineering, Mechanical Engineering, Electrical & Computer Engineering, and School of Civil & Environmental Engineering, Cornell University, Ithaca, NY, USA
In this paper, we propose a hierarchical multi-agent reinforcement learning framework for coordinating robot teams in warehouse environments under communication loss. We partition the robot team into groups, with centralized coordination within each group and distributed coordination across groups. Each group uses a recurrent predictor to estimate unavailable interaction information due to communication loss. A higher-level policy then generates a compact coordination reference that conditions the local control policy within each group. A predictive safety filter evaluates and modifies the proposed controls when they violate safety constraints. Simulation results show improved task completion under communication loss, reduced communication growth as the team size increases, and safe operation in the tested scenarios.
Figures & tables
Fig. 1: Overview of the proposed hierarchical coordination framework.
Fig. 2: Training convergence of the proposed method and the comparison baseline under different robot team sizes, with a 0.2 communication loss rate.
Fig. 3: Success rate comparison between the proposed method and IC3Net for teams of 40 and 60 robots under communication loss probabilities p∈{0.2,0.4,0.6} .
Fig. 4: (a). Communication load per time step for the proposed method and the fully decentralized baseline. (b). Ratio between payloads of the two methods and the analytical result.
Robust multi-agent coordination relies heavily on inter-agent communication, which is frequently disrupted by physical and environmental constraints in real-world deployments. To maintain operation during these intermittent communication failures, agents can employ internal prediction models to estimate missing shared state information. However, predictors trained with standard reconstruction objectives treat all transitions equally. In a Reinforcement Learning context, this forces the model to waste capacity learning stochastic exploration noise and the outdated dynamics of suboptimal policies. In this paper, we propose a value-aware extension of Multi-Agent Observation Sharing under Communication Dropout (MARO) to patch communication gaps; we refer to this method as Value-Aware MARO. By dynamically weighting the predictor's loss function using advantage estimates derived from the underlying actor-critic architecture, our objective explicitly couples the predictor's learning process to the policy's evolution. This formulation focuses the model's capacity on the intentional, high-return dynamics actively reinforced by the agents. We evaluate our framework on several tasks within the Multi-Agent Particle Environment under varying communication reliability levels. Experimental results demonstrate that our approach maintains performance under declining communication reliability, particularly below 40%. While our method performs comparably in tasks where the baseline already maintains high coordination, our value-aware weighting effectively prevents the performance collapse observed in the standard predictor during high-attrition scenarios. In these environments, our method achieves an average improvement in mean returns of more than 20% and reduces performance variance by a mean of 64.7% compared to the standard unweighted baseline.
Kemal Devrim Kafadar, Eren Özaltun, Mahmud Efnan Şanlı +4
Department of Computer Engineering, Istanbul Technical University, Istanbul, Türkiye. · Department of Computer Science, University of Stuttgart, Stuttgart, Germany. · Istanbul Technical University Artificial Intelligence and Data Science Application and Research Center, Istanbul, Türkiye. +1
Emergent communication enables partially observant Autonomous Mobile Robots (AMRs) to coordinate effectively in decentralized multi-agent reinforcement learning (MARL) settings. However, existing approaches often struggle with unstable communication protocols, ungrounded message semantics, and interference between communication learning and policy optimization, leading to degraded coordination over time. We propose SCALE-COMM (Shared, Contrastively-Aligned Latent Embeddings for COMMunication), a self-supervised framework for learning compact, stable, and policy-relevant communication representations. SCALE-COMM decouples communication learning from policy optimization by training low-dimensional latent messages that capture task-relevant planning and traffic information, while enforcing consistency across agents and time. Across standard MARL benchmarks and a realistic warehouse coordination task, SCALE-COMM consistently outperforms existing communication frameworks in both representation quality and task performance. The learned communication space yields improved stability, sample efficiency, and throughput under policy fine-tuning, demonstrating the effectiveness of representation-driven communication for scalable multi-agent coordination.
Mahmoud Abouelyazid, Eman Hammad
Electrical and Computer Engineering Department Texas A&M University, College Station, TX, USA
Cooperative multi-agent reinforcement learning (MARL) often relies on communication to mitigate partial observability, yet most existing protocols treat messages as flat dense vectors detached from the structure of the observations they summarize. This design overlooks an important source of inductive bias in many cooperative environments, where observations naturally follow a hierarchy such as groups and entities. We propose \textsc{HiComm}, a plug-in communication module that grounds messages in the sender's hierarchical observation. \textsc{HiComm} is receiver-driven: the receiver issues a query, and the hierarchy is resolved through a three-stage decoding process that first selects a group, then a sender, and then an entity within that group, returning the corresponding feature slice as the message. This converts communication from unstructured vector transmission into structured information retrieval over the sender's observation hierarchy. We instantiate this mechanism with Straight-Through Gumbel-Softmax for differentiable discrete selection and a lightweight shared projection design that attaches to standard MARL pipelines. Experiments across cooperative MARL tasks with different observation structures and coordination demands show that \textsc{HiComm} matches or outperforms representative learned communication baselines while reducing communication volume by up to 23× per receiver per episode.
Runze Zhao, Dongruo Zhou, Sumit Kumar Jha +2
Luddy School of Informatics, Computing, and Engineering Indiana University Bloomington Bloomington, IN 47408 · Department of Computer & Information Science & Engineering University of Florida Gainesville, FL 32611 · Department of Electrical Engineering and Computer Science United States Military Academy West Point, NY 10996 +1