Efficient cooperation is challenging due to the usual partial observability of each agent in multi-agent reinforcement learning. Recurrent networks encode local interaction histories, but their hidden representations provide limited insight into the information underlying individual decisions. To address these challenges, we propose a novel interpretable framework, called escaping local views (ELV), which introduces semantically structured latent concepts to render policy decisions transparent. Specifically, each agent extracts low-dimensional semantic concepts from its local observation and action-observation trajectory. These concepts are jointly encoded into a contextual latent variable via a variational autoencoder (VAE), which builds a bridge between local views and global semantics. To explicitly model the decision of each agent, we employ a dual-path attention mechanism in which one module estimates the salience of individual concepts relative to the global context, while the other captures higher-order cooperative patterns with pairwise concept interactions. Furthermore, we incorporate a concept prediction module that derives an intrinsic reward from next-concept prediction errors, which incentivizes agents to explore regions of semantic novelty. Experiments in multiple environments verify that ELV not only achieves competitive performance but also explicitly provides how agents reason about their decisions.
Figures & tables
Figure 1: At each timestep t , agent i encodes its local history τit into a hidden state hit , from which ELV extracts concept embeddings {ci,kt}k=1K and keys {ki,k} via shared projections. The concepts of all agents are aggregated and passed through a VAE to infer a global context z , which yields a query vector q . First-order weights αi,k are obtained by softmax-normalizing dot-product scores q⊤ki,k . To improve coordination, a second-order module scores concept pairs (ci,k,cj,m) ( j=i ) against q , aggregates over other agents to form βi,k , and normalizes across k . Finally, ELV reweights concepts with αi,k+βi,k and takes a linear head to sum the individual value Qi(τit,uit) .
Figure 2: Performance comparison on two tasks of LBF.
Figure 3: Performance comparison on hard and super-hard scenarios.
Figure 4: Ablation of intrinsic reward and concept decoder for ELV.
Figure 5: Ablation of the concept interaction module for ELV.
Figure 6: Ablation study with different numbers of concepts for ELV.
Figure 7: The t-SNE visualization of learned concept embeddings.
Figure 8: Dependency strength over health and concept assignment heatmap.
Figure 9: Visualization of agent collaboration dynamics.
Appendix figures & tables14 assets
Supplementary material from the paper’s appendix.
Appendix
Figure 10: Three benchmarks used in our experiments.
Hyperparameter
Value
Description
Max Player Level
3
Maximum Agent Level Attribute
Max Episode Length
50
Maximum Timesteps Per Episode
Batch Size
32
Number of Episodes Per Update
Test Interval
10,000
Frequency of Evaluating Performance
Test Episodes
32
Number of Episodes to Test
Replay Batch Size
5000
Maximum Number of Episodes Stored in Memory
Appendix
Table 1: Experimental Settings of Level-Based Foraging.
Hyperparameter
Value
Description
Difficulty
7
Enemy Units with Built-in AI Difficulty
Batch Size (SMAC)
32
Number of Episodes Per Update
Batch Size (SMACv2)
128
Number of Episodes Per Update
Test Interval
10,000
Frequency of Evaluating Performance
Test Episodes
32
Number of Episodes to Test
Replay Batch Size
5000
Maximum Number of Episodes in Memory
Appendix
Table 2: Experimental settings of SMAC and SMACv2.
Map Name
Ally Units
Enemy Units
Timesteps
Scenario Type
2s3z
2 Stalkers & 3 Zealots
2 Stalkers & 3 Zealots
2M
Easy
3s5z
3 Stalkers & 5 Zealots
3 Stalkers & 5 Zealots
2M
Easy
1c3s5z
1 Colossus , 3 Stalkers & 5 Zealots
1 Colossus , 3 Stalkers & 5 Zealots
2M
Easy
2s_vs_1sc
2 Stalkers
1 Spine Crawler
2M
Easy
5m_vs_6m
5 Marines
6 Marines
2M
Hard
8m_vs_9m
8 Marines
9 Marines
2M
Hard
Appendix
Table 3: The detailed specifications of SMAC scenarios used in our evaluation.
Scenario Name
Allies
Enemies
Timesteps
Unit Composition (Randomized)
Terran_5_vs_5
5
5
1M
Marines, Marauders, Medivacs
Terran_10_vs_10
10
10
2M
Zerg_5_vs_5
5
5
1M
Zerglings, Hydralisks, Banelings
Zerg_10_vs_10
10
10
2M
Protoss_5_vs_5
5
5
1M
Stalkers, Zealots, Colossi
Protoss_10_vs_10
10
10
2M
Appendix
Table 4: The SMACv2 scenarios used in our evaluation.
Component
Hyper-parameters
Value
Concept Number ( K )
16
Concept
Concept Dimension ( d )
16
Extraction
History Hidden Dim
64
Latent Context Dim ( z )
32
Variational
VAE Loss Weight λ1
0.1
Autoencoder
Rec. Loss Weight λ2
0.1
Appendix
Table 5: Hyper-parameters of ELV.
Figure 11: Median test win rate % on ELV with VDN mixing network.
Figure 12: Test win rate % for six extra scenarios of SMAC benchmark.
Figure 13: Median test win rate % on SMACv2.
Figure 14: Visualization of exploration coverage via spatial heatmaps.
Concept ID
Dominant feature
Semantic label
C14
move
Movement direction
C2
enemy1
Enemy unit 1 tracking
C13
enemy2
Enemy unit 2 tracking
C9
enemy3
Enemy unit 3 tracking
C5
enemy4
Enemy unit 4 tracking
C0
enemy5
Enemy unit 5 tracking
Appendix
Table 6: Attribution-based semantic assignments of learned concept slots on 3s_vs_5z . Concept IDs refer to slots in the analyzed model.
Concept ID
Win rate ( p=0 )
Win rate ( p=1 )
C9
0.00
0.30
C11
0.00
0.60
C14
0.10
0.70
C2
0.10
0.20
C13
0.60
0.60
C4
0.60
0.90
Appendix
Table 7: Test-time concept interventions on 3s_vs_5z . Each setting uses 10 evaluation episodes. The unmodified policy’s win rate is 1.00 ; all values are proportions.
Figure 15: Temporal validation of the assignment of Agent 0’s Concept 2 to enemy unit 1. The panels show agent–enemy distance with enemy visibility (top), concept activation (middle), and enemy health (bottom). The dotted vertical line marks the enemy death event at t=64 .
Figure 16: Mean test return on LBF with 6 players and 4 foods on a 10×10 grid, using observation radii of 2, 4, and 8. ELV and MA2E use the same QMIX mixing network.
Communication is a key component in multi-agent reinforcement learning (MARL) for mitigating partial observability, yet prior approaches often rely on inefficient information exchange or fail to transmit sufficient state information. To address this, we propose LLM-driven Multi-Agent Communication (LMAC), which leverages an LLM's reasoning capability to design a communication protocol that enables all agents to reconstruct the underlying state as accurately and uniformly as possible. LMAC iteratively refines the protocol using an explicit state-awareness criterion, improving state recovery while narrowing differences in agents' knowledge. Experiments on diverse MARL benchmarks show that LMAC improves state reconstruction across agents and yields substantial performance gains over prior communication baselines.
Sangjun Bae, Yisak Park, Sanghyeon Lee +1
Graduate School of Artificial Intelligence, UNIST, Ulsan, South Korea.
To promote cooperation in Multi-Agent Reinforcement Learning, the reward signals of all agents can be aggregated together, forming global rewards that are commonly known as the fully cooperative setting. However, global rewards are usually noisy because they contain the contributions of all agents, which have to be resolved in the credit assignment process. On the other hand, using local reward benefits from faster learning due to the separation of agents' contributions, but can be suboptimal as agents myopically optimize their own reward while disregarding the global optimality. In this work, we propose a method that combines the merits of both approaches. By using a graph of interaction between agents, our method discerns the individual agent contribution in a more fine-grained manner than a global reward, while alleviating the cooperation problem with agents' local reward. We also introduce a practical approach for approximating such a graph. Our experiments demonstrate the flexibility of the approach, enabling improvements over the traditional local and global reward settings.
Bang Giang Le, Viet Cuong Ta
Human-Machine Interaction Laboratory VNU University of Engineering and Technology, Hanoi, Vietnam
Decentralized Multi-Agent Reinforcement Learning (MARL) methods allow for learning scalable multi-agent policies, but suffer from partial observability and induced non-stationarity. These challenges can be addressed by introducing mechanisms that facilitate coordination and high-level planning. Specifically, coordination and temporal abstraction can be achieved through communication (e.g., message passing) and Hierarchical Reinforcement Learning (HRL) approaches to decision-making. However, optimization issues limit the applicability of hierarchical policies to multi-agent systems. As such, the combination of these approaches has not been fully explored. To fill this void, we propose a novel and effective methodology for learning multi-agent hierarchies of message-passing policies. We adopt the feudal HRL framework and rely on a hierarchical graph structure for planning and coordination among agents. Agents at lower levels in the hierarchy receive goals from the upper levels and exchange messages with neighboring agents at the same level. To learn hierarchical multi-agent policies, we design a novel reward-assignment method based on training the lower-level policies to maximize the advantage function associated with the upper levels. Results on relevant benchmarks show that our method performs favorably compared to the state of the art.
Tommaso Marzi, Cesare Alippi, Andrea Cini
IDSIA USI-SUPSI, Università della Svizzera italiana, Lugano, Switzerland · Politecnico di Milano, Milan, Italy · EPFL, Lausanne, Switzerland