Independent Multi-Agent Reinforcement Learning with Counterfactual Semantic-Social World Models
Authors: Fernando Martinez, Tao Li, Yingdong Lu, Juntao Chen
Organizations: Department of Computer and Information Sciences, Fordham University, New York, NY, 10023, USA · Department of Systems Engineering, City University of Hong Kong, Hong Kong SAR, 999077, China · IBM Research, Yorktown Heights, NY, 10598, USA
Fully decentralized multi-agent reinforcement learning (MARL), also referred to as independent learning, requires each agent to learn and act using only its local information and experience, without a centralized critic or inter-agent communication. Such a stringent information structure renders the conventional reward signal ambiguous. A poor return may result from an ineffective ego action, an incompatible teammate response, or an effective opponent response, yet scalar rewards alone do not reveal which explanation is responsible. We argue that agents can learn more effectively by prospectively comparing the consequences of candidate actions rather than diagnosing failures only from realized returns. We introduce CASTLE (Counterfactual Action-conditioned Semantic Tokens for Local Execution in Decentralized MARL), an offline-training, online-in-context guidance framework with two complementary world models. A Local Dynamics World Model, offline pre-trained over agents' local trajectories, summarizes the agent's local trajectory dynamics and partial observability, while a Semantic-Social World Model predicts compact short-horizon task and social consequences for each candidate ego action. The latter is trained from counterfactual simulator rollouts that expose plausible teammate and opponent responses to alternative actions taken from the same logged rollout state. During online learning and execution, both world models remain frozen and are queried by agents using only locally available information. Their prediction logits provide in-context guidance to an independent PPO policy. Across 30 matched seeds on Tag, Spread, and Adversary in the benchmark multi-particle environments, our proposed CASTLE achieves the highest mean final score among the evaluated methods, exceeding the strongest baseline on each task by 10.67, 6.46, and 0.33 normalized points, respectively.
Figures & tables
Fig. 1: Architecture of CASTLE . Offline pretraining produces frozen LDWM checkpoints and an SSWM that emits one consequence token zt,ki for each candidate ego action. During decentralized IPPO, a trainable actor combines base policy logits with an annealed gated utility branch computed from the frozen SSWM tokens. Online gradients update only the orange PPO components; blue world-model modules remain frozen.
Fig. 2: Target-task predictive accuracy over the SSWM rollout horizon. Solid lines indicate target-included pretraining and dashed lines held-out pretraining, evaluated on identical target-task histories. Correlations are computed across candidate actions; shading denotes 95% confidence intervals.
Method
tag
spread
adversary
Total Average
Decentralized baselines
IPPO
78.07±8.67
70.69±12.29
100.34±0.92
83.03
IPPO + Transformer
70.83±10.56
79.64±5.58
102.91±1.41
84.46
IQL
76.75±7.48
76.41±9.24
95.55±1.84
82.90
IQL + Transformer
65.91±8.04
86.11±4.33
101.05±1.33
84.36
DPO
74.46±7.31
82.03±4.47
98.77±1.22
85.09
TABLE I: D4RL-normalized final performance over 30 matched seeds, reported as mean ± standard deviation. Higher is better. In-domain (R) adds target-task random trajectories to the held-out source pool, while in-domain (R+M) adds target-task random and intermediate trajectories. Bold and underlined entries are the best and second-best means in each column.
Fig. 3: Raw online returns over 30 matched seeds. Lines show means and shaded regions 95% confidence intervals. Curves use a 25-update moving average.
Fig. 4: Raw online returns over 30 matched seeds. Lines show means and shaded regions 95% confidence intervals. Curves use a 25-update moving average.
Fig. 5: Effect of consequence guidance over 30 matched seeds. Top : KL divergence between the guided policy and the base IPPO actor. Bottom : correlation between the guidance-induced change in the sampled action’s log probability and its realized local GAE advantage.
Parameter
Value
Local Dynamics World Model (LDWM)
Deterministic / hidden dimension
256 / 256
Stochastic variables / categories
32 / 32
Context dimension / uniform mixture
1280 / 0.01
Adam learning rate / epsilon
3×10−4 / 10−5
Gradient norm limit / KL free nats
100 / 0.5
TABLE II: Model architecture and LDWM reference pretraining.