Independent Multi-Agent Reinforcement Learning with Counterfactual Semantic-Social World Models
Authors: Fernando Martinez, Tao Li, Yingdong Lu, Juntao Chen
Organizations: Department of Computer and Information Sciences, Fordham University, New York, NY, 10023, USA · Department of Systems Engineering, City University of Hong Kong, Hong Kong SAR, 999077, China · IBM Research, Yorktown Heights, NY, 10598, USA
Fully decentralized multi-agent reinforcement learning (MARL), also referred to as independent learning, requires each agent to learn and act using only its local information and experience, without a centralized critic or inter-agent communication. Such a stringent information structure renders the conventional reward signal ambiguous. A poor return may result from an ineffective ego action, an incompatible teammate response, or an effective opponent response, yet scalar rewards alone do not reveal which explanation is responsible. We argue that agents can learn more effectively by prospectively comparing the consequences of candidate actions rather than diagnosing failures only from realized returns. We introduce CASTLE (Counterfactual Action-conditioned Semantic Tokens for Local Execution in Decentralized MARL), an offline-training, online-in-context guidance framework with two complementary world models. A Local Dynamics World Model, offline pre-trained over agents' local trajectories, summarizes the agent's local trajectory dynamics and partial observability, while a Semantic-Social World Model predicts compact short-horizon task and social consequences for each candidate ego action. The latter is trained from counterfactual simulator rollouts that expose plausible teammate and opponent responses to alternative actions taken from the same logged rollout state. During online learning and execution, both world models remain frozen and are queried by agents using only locally available information. Their prediction logits provide in-context guidance to an independent PPO policy. Across 30 matched seeds on Tag, Spread, and Adversary in the benchmark multi-particle environments, our proposed CASTLE achieves the highest mean final score among the evaluated methods, exceeding the strongest baseline on each task by 10.67, 6.46, and 0.33 normalized points, respectively.
Figures & tables
Fig. 1: Architecture of CASTLE . Offline pretraining produces frozen LDWM checkpoints and an SSWM that emits one consequence token zt,ki for each candidate ego action. During decentralized IPPO, a trainable actor combines base policy logits with an annealed gated utility branch computed from the frozen SSWM tokens. Online gradients update only the orange PPO components; blue world-model modules remain frozen.
Fig. 2: Target-task predictive accuracy over the SSWM rollout horizon. Solid lines indicate target-included pretraining and dashed lines held-out pretraining, evaluated on identical target-task histories. Correlations are computed across candidate actions; shading denotes 95% confidence intervals.
Method
tag
spread
adversary
Total Average
Decentralized baselines
IPPO
78.07±8.67
70.69±12.29
100.34±0.92
83.03
IPPO + Transformer
70.83±10.56
79.64±5.58
102.91±1.41
84.46
IQL
76.75±7.48
76.41±9.24
95.55±1.84
82.90
IQL + Transformer
65.91±8.04
86.11±4.33
101.05±1.33
84.36
DPO
74.46±7.31
82.03±4.47
98.77±1.22
85.09
TABLE I: D4RL-normalized final performance over 30 matched seeds, reported as mean ± standard deviation. Higher is better. In-domain (R) adds target-task random trajectories to the held-out source pool, while in-domain (R+M) adds target-task random and intermediate trajectories. Bold and underlined entries are the best and second-best means in each column.
Fig. 3: Raw online returns over 30 matched seeds. Lines show means and shaded regions 95% confidence intervals. Curves use a 25-update moving average.
Fig. 4: Raw online returns over 30 matched seeds. Lines show means and shaded regions 95% confidence intervals. Curves use a 25-update moving average.
Fig. 5: Effect of consequence guidance over 30 matched seeds. Top : KL divergence between the guided policy and the base IPPO actor. Bottom : correlation between the guidance-induced change in the sampled action’s log probability and its realized local GAE advantage.
Parameter
Value
Local Dynamics World Model (LDWM)
Deterministic / hidden dimension
256 / 256
Stochastic variables / categories
32 / 32
Context dimension / uniform mixture
1280 / 0.01
Adam learning rate / epsilon
3×10−4 / 10−5
Gradient norm limit / KL free nats
100 / 0.5
TABLE II: Model architecture and LDWM reference pretraining.
In multi-agent reinforcement learning (MARL), inter-agent communication is effective for improving performance under partial observability. Representation learning-based approaches enable decentralized agents to learn messages grounded in their own observations, but they rely only on current observations and cannot convey information accumulated over time. We propose Dreamer-CPC, a decentralized model-based MARL method that integrates message learning based on Collective Predictive Coding (CPC) into the world model of DreamerV3. Each agent independently maintains a world model and a message module, and infers and exchanges messages from the latent states of the world model that reflect the history of past observations and actions. We evaluated Dreamer-CPC in two environments: Observer, a non-cooperative information-sharing task, and CatchApple, a newly introduced task in which task-relevant observations are temporarily missing. In both environments, Dreamer-CPC outperformed IPPO-CPC, an existing CPC-based method that generates messages from current observations, as well as no-communication baselines. In particular, in CatchApple, Dreamer-CPC achieved 4 to 5 times the episode return of IPPO-CPC, demonstrating effective coordination where other methods fail due to missing observations. These results suggest that communication grounded in the latent dynamics of world models can support decentralized decision-making when current observations alone are insufficient.
World models improve sample efficiency by training policies on imagined trajectories, but their usefulness depends on learning representations that capture the information needed for future control. We study whether self-supervised joint-embedding prediction (JEPA) can provide this learning signal for multi-agent reinforcement learning. We introduce MA-JEPA, a stochastic world model that replaces observation reconstruction with prediction of target representations, enabling model-based multi-agent reinforcement learning with centralized training and decentralized execution. A categorical latent state and a causal Transformer are trained with posterior and action-conditioned dynamics prediction objectives and are then used for actor-critic learning from latent imagination. A training-only joint predictor conditions on all agents' local states and actions to predict each agent's next local observation embedding. These predictions are passed through the same local posterior used during real interaction with a centralized critic that is used only for value learning, with execution remaining decentralized. Our experiments show that this architecture performs strongly on SMAC, matching or exceeding the strongest reported comparator mean win rate on four of eight evaluated maps.
Brandon Gary Kaplowitz, Osaze James Obahor, Christian Schroeder de Witt
Department of Engineering Science, University of Oxford
A key challenge in multi-agent reinforcement learning (MARL) lies in designing learning signals that effectively promote coordination among agents. Designing such signals requires estimating how one agent's current action affects its teammates over future interaction steps. To address this, we introduce Multi-step Advantage-Gated Interventional Causal MARL (MAGIC), a framework that estimates multi-step action effects between agents and selectively converts them into intrinsic rewards. MAGIC uses counterfactual action interventions to compare teammate futures under factual and counterfactual branches, and introduces a gate based on advantage to direct exploration toward beneficial behaviors aligned with the task goal. Experiments on Multi-Agent Particle Environments (MPE) and StarCraft micromanagement benchmarks (SMAC and SMACv2) show that MAGIC consistently outperforms leading prior methods, with average relative final performance improvements of 26.9% and 10.1%, respectively.
Haohan Yu, Jinmiao Cong, Shengzhi Wang +2
1Dalian University of Technology · 2Microsoft, Beijing, China