cs.LGOct 7, 2026

Continual Graph Multi-Agent Reinforcement Learning

Authors: Tommaso Marzi, Ahmed Hendawy, Jan Peters, Carlo D'Eramo, Andrea Cini, Cesare Alippi

Organizations: IDSIA USI-SUPSI, Università della Svizzera italiana, Lugano, Switzerland · Technical University of Darmstadt, Darmstadt, Germany · Robotics Institute Germany (RIG) · German Research Center for AI (DFKI) · University of Würzburg, Würzburg, Germany · EPFL, Lausanne, Switzerland · Politecnico di Milano, Milan, Italy

Abstract

In Continual Multi-Agent Reinforcement Learning (CMARL), agents learn cooperative policies across sequences of tasks, aiming to adapt effectively to new tasks while preserving the ability to solve previously encountered ones. In many applications, tasks differ in their underlying structure, which can represent, for example, distinct operational conditions or target configurations (e.g., different network topologies in power grids or arrangements in formation control). Existing CMARL methods lack dedicated mechanisms to leverage this structural information when learning new tasks, failing to promote transfer and mitigate forgetting. To fill this gap, we propose Continual Graph Multi-Agent Reinforcement Learning (CGMARL), a novel framework for CMARL problems in which task sequences are mapped into a series of attributed graphs, each modeling a task-specific structure. In CGMARL, each graph determines the environment dynamics (next states and/or rewards) and the number of agents for the corresponding task. Then, we present Graph-based Formation (GRAFO), the first CGMARL benchmark, and show how forgetting arises in this setting. Finally, to address this limitation, we propose Frozen Graph Encoder (FROG), a method that relies on a frozen graph backbone to preserve past structural information in graph-based CMARL policies. Experiments on GRAFO show that pairing FROG with existing CL methods substantially improves performance on multiple CGMARL scenarios.

Figures & tables

Appendix figures & tables7 assets

Supplementary material from the paper’s appendix.

Appendix

Explore similar work

CardsList
  1. Coordination Graphs for Constrained Multi-Agent Reinforcement Learning

    Jun 1, 2026Santiago Amaya-Corredor, Miguel Calvo-Fullana, Anders JonssonMulti-Agent Reinforcement LearningCoordination

  2. GCT-MARL: Graph-Based Contrastive Transfer for Sample-Efficient Cooperative Multi-Agent Reinforcement Learning

    Jun 23, 2026Animesh Animesh, Satheesh K Perepu, Kaushik DeyMulti-Agent Reinforcement LearningIterative Co-Training

  3. A Self-Evolving Default Action for Cooperative Tasks with Continuous Action Space

    Jul 21, 2026Shuangyao HuangMulti-Agent Reinforcement LearningAction Space