Authors: Tommaso Marzi, Ahmed Hendawy, Jan Peters, Carlo D'Eramo, Andrea Cini, Cesare Alippi
Organizations: IDSIA USI-SUPSI, Università della Svizzera italiana, Lugano, Switzerland · Technical University of Darmstadt, Darmstadt, Germany · Robotics Institute Germany (RIG) · German Research Center for AI (DFKI) · University of Würzburg, Würzburg, Germany · EPFL, Lausanne, Switzerland · Politecnico di Milano, Milan, Italy
In Continual Multi-Agent Reinforcement Learning (CMARL), agents learn cooperative policies across sequences of tasks, aiming to adapt effectively to new tasks while preserving the ability to solve previously encountered ones. In many applications, tasks differ in their underlying structure, which can represent, for example, distinct operational conditions or target configurations (e.g., different network topologies in power grids or arrangements in formation control). Existing CMARL methods lack dedicated mechanisms to leverage this structural information when learning new tasks, failing to promote transfer and mitigate forgetting. To fill this gap, we propose Continual Graph Multi-Agent Reinforcement Learning (CGMARL), a novel framework for CMARL problems in which task sequences are mapped into a series of attributed graphs, each modeling a task-specific structure. In CGMARL, each graph determines the environment dynamics (next states and/or rewards) and the number of agents for the corresponding task. Then, we present Graph-based Formation (GRAFO), the first CGMARL benchmark, and show how forgetting arises in this setting. Finally, to address this limitation, we propose Frozen Graph Encoder (FROG), a method that relies on a frozen graph backbone to preserve past structural information in graph-based CMARL policies. Experiments on GRAFO show that pairing FROG with existing CL methods substantially improves performance on multiple CGMARL scenarios.
Figures & tables
Figure 1: Scheme of the CGMARL problem. In the left panel, we illustrate the general setup, in which edges are generated dynamically across timesteps using task-specific connectivity functions (indicated by different colors for edges). In the right panel, we illustrate relevant classes of the CGMARL problem. In Topology Switch (TS), both nodes and edges can change across tasks. In Connectivity Switch (CS), edges can change while the node set is fixed. In Graph Features Switch (GFS), the topology is kept static, and graph features can change; this is highlighted by using different dashed patterns to represent nodes and/or edges.
Figure 2: Example of a GRAFO task with five agents (circles). Left: agent graph identifying the task. Right: current state of the environment, where the green solid cross is the landmark known to all agents, i.e., the target position of the green agent, while the dashed crosses are the target positions of other agents. These positions are defined based on the graph, but are not explicitly provided.
Table 1: Solved percentage at test time (median ± standard deviation of 5 seeds) for two GRAFO tasks in both zero-shot transfer (TRAN.) and continual (CONT.) settings.
Figure 3: Boxplot showing the distribution of 10 independent seeds for the different CL methods (with and without FROG) across the three scenarios, i.e., TS (top), CS (middle), and GFS (bottom). Boxes represent the IQR, horizontal lines indicate medians, and whiskers extend to 1.5×IQR .
AFP ↑
AF ↑
AP ↑
FROG
0.69±0.13
−0.33±0.14
0.99±0.02
FROG (RND)
0.38±0.037
−0.29±0.04
0.64±0.03
Table 2: Aggregated medians (average ± standard deviation) for the different CL methods combined with FROG and FROG (RND) on the TS sequence. Complete boxplots are reported in Fig. 7 .
Figure 4: Boxplot showing the distribution of 10 independent seeds for the different CL methods (with and without FROG) across the TS (inc) and TS (dec) sequences, i.e., the incremental/decremental variants of the TS sequence (top and bottom panels, respectively). Boxes represent the IQR, horizontal lines indicate medians, and whiskers extend to 1.5×IQR .
Appendix figures & tables7 assets
Supplementary material from the paper’s appendix.
Appendix
Figure 5: One-step transition in GRAFO. At the beginning of the episode, an initial state in which the agents (colored circles) and known landmark (solid, green cross) positions are randomly assigned is generated. Then, at each step, agents take actions depending on the current state. Each agent’s reward is the negative of the distance to its corresponding target (dashed cross with the same color). In the case of two symmetric agents, i.e., with a non-empty set K of structurally-equivalent nodes, a linear sum assignment problem is solved to assign the targets.
GRAFO- 3
GRAFO- 4
GRAFO- 5
GRAFO- 6 (1)
GRAFO- 6 (3)
GRAFO- 7
InforMARL
1±0
1±0
1±0
0.994±0.009
1±0
0.997±0.007
InforMARL (no GNN)
0.531±0.013
0.328±0.091
0.297±0.258
0±0
0.172±0.169
0±0
Appendix
Table 3: Solved percentage of InforMARL and the variant without GNN in the actor on the single tasks of the TS sequence (reported as median ± standard deviation of 5 seeds).
Figure 6: Boxplot showing the distribution of 10 independent seeds in the sensitivity analysis of λμ for EWC (M) in the TS sequence. Specifically, we consider different values for the regularization coefficient λμ of the action prediction module. The regularization coefficient of the graph encoder is fixed to λψ=1010 . When the coefficient for the action module is λμ=1010 (pink box), the results coincide with those reported in Fig. 3 . We also report the results achieved by FROG, where we remark that the graph encoder is kept fixed, i.e., the coefficient λμ=104 regularizes only the action prediction module. Boxes represent the IQR, horizontal lines indicate medians, and whiskers extend to 1.5×IQR .
Figure 7: Boxplot showing the distribution of 10 independent seeds for FROG and FROG (RND) combined with different CL methods in the TS scenario. Boxes represent the IQR, horizontal lines indicate medians, and whiskers extend to 1.5×IQR .
Figure 8: Boxplot showing the distribution of 10 independent seeds for different CL methods in the TS scenario with and without conditioning on task identities. Boxes represent the IQR, horizontal lines indicate medians, and whiskers extend to 1.5×IQR .
Figure 9: GRAFO tasks used to define the CGMARL sequences. The green circle is the reference node, i.e., the agent for which the target position corresponds to the reference landmark. Blue circles represent agents for which the corresponding target positions g=(g1,…,gn) are unknown, and are reported as relative shifts (one per agent) with respect to the green node. Dashed blue nodes indicate agents in K⊂G with a structurally equivalent role in the graph, i.e., that can be exchanged without altering the topology and for which a linear assignment with respect to the target positions is solved.
Context
Parameter
Value
Architecture (InforMARL)
Activation
ReLU
Num. of hidden layers (action module)
2
Hidden feature dim (action module)
128
Recurrent unit (action module)
False
Appendix
Table 4: Hyperparameters used in our experiments. We note that the parameters specific to the online variants of EWC and MAS (with and without FROG), i.e., online decay, importance episodes, importance steps, and normalization, are common for all methods.
Constrained Multi-agent reinforcement learning (CMARL) faces two intertwined challenges: the joint action space grows exponentially with the number of agents, and additional requirements couple agents in ways that reward structure alone does not capture. We introduce Coordination Graphs for Constrained Multi-Agent Reinforcement Learning (CG-CMARL), a framework that addresses both challenges by combining coordination graphs with Lagrangian duality. The system decomposes the joint problem into pairwise regions, each served by a set of shared Q-functions, one for the primary objective and one for each of the constraints, so that the number of learned models is independent of the number of agents. At execution time, Max-Sum message passing coordinates actions across the factor graph, while a Lagrangian multiplier controls the objective--constraint tradeoff, allowing a single trained model to trace a Pareto front without retraining. We provide convergence guarantees under mild conditions, together with a compositional error bound that decomposes into separate interpretable sources, each traceable to a specific design choice and independently controllable. Experiments on cooperative navigation tasks (where teams of up to 10 agents must coordinate to reach target positions while satisfying pairwise constraints) show that our method produces Pareto fronts dominating established baselines trained at fixed reward-shaping ratios, while scaling to team sizes where centralized approaches become intractable.
Santiago Amaya-Corredor, Miguel Calvo-Fullana, Anders Jonsson
Department of Engineering, Universitat Pompeu Fabra, Barcelona, Spain
In cooperative multi-agent reinforcement learning (MARL), from a deployment perspective, it is challenging and expensive to train agents from scratch for each new environment or task. In this work, we propose GCT-MARL, a transfer learning framework that builds on the multi-view graph contrastive backbone of MAIL and augments it with a per-view, adaptively weighted alignment loss and a two-phase training protocol specifically designed for transfer across populations of varying sizes and compositions. We empirically demonstrate that the proposed framework markedly accelerates convergence on the target task relative to from-scratch training, in both homogeneous (within-faction, varying N) and heterogeneous (cross-faction and mixed unit-type) transfer scenarios. Furthermore, we show that the framework naturally supports continual learning by sequentially chaining the two-phase transfer protocol across a series of related tasks. Overall, this work provides a unified approach to mitigating key limitations in current MARL transfer methods with new insights at both methodological and empirical levels.
Animesh Animesh, Satheesh K Perepu, Kaushik Dey
Department of Artificial Intelligence, Indian Institute of Technology Kharagpur, India · Work done during research internship at Ericsson Research, Bangalore, India · Ericsson Research, Bangalore, India
Counterfactual credit assignment has proven effective in multi-agent reinforcement learning (MARL) for discrete action spaces, yet its extension to continuous-action cooperative tasks remains challenging. Existing methods that approximate the counterfactual baseline via Monte Carlo sampling often introduce bias into policy gradients and fail to guarantee convergence to local optima, as the sampled actions may not have been sufficiently trained. To address these limitations, we propose SAFE, a novel MARL framework that employs a counterfactual baseline conditioned on a self-evolving default action sampled from each agent's experience buffer. This design naturally extends to continuous action spaces without relying on additional simulations, reward models, or environment-specific prior knowledge. The baseline accurately quantifies each agent's contribution, and introduces no bias into the deterministic policy gradient, ensuring convergence to local optima. Extensive experiments on cooperative vehicular tasks demonstrate that SAFE consistently outperforms state-of-the-art models.
Shuangyao Huang
Xi’an Jiaotong-Liverpool University School of Internet of Things