cs.MAOct 6, 2026

Independent Multi-Agent Reinforcement Learning with Counterfactual Semantic-Social World Models

Authors: Fernando Martinez, Tao Li, Yingdong Lu, Juntao Chen

Organizations: Department of Computer and Information Sciences, Fordham University, New York, NY, 10023, USA · Department of Systems Engineering, City University of Hong Kong, Hong Kong SAR, 999077, China · IBM Research, Yorktown Heights, NY, 10598, USA

Abstract

Fully decentralized multi-agent reinforcement learning (MARL), also referred to as independent learning, requires each agent to learn and act using only its local information and experience, without a centralized critic or inter-agent communication. Such a stringent information structure renders the conventional reward signal ambiguous. A poor return may result from an ineffective ego action, an incompatible teammate response, or an effective opponent response, yet scalar rewards alone do not reveal which explanation is responsible. We argue that agents can learn more effectively by prospectively comparing the consequences of candidate actions rather than diagnosing failures only from realized returns. We introduce CASTLE (Counterfactual Action-conditioned Semantic Tokens for Local Execution in Decentralized MARL), an offline-training, online-in-context guidance framework with two complementary world models. A Local Dynamics World Model, offline pre-trained over agents' local trajectories, summarizes the agent's local trajectory dynamics and partial observability, while a Semantic-Social World Model predicts compact short-horizon task and social consequences for each candidate ego action. The latter is trained from counterfactual simulator rollouts that expose plausible teammate and opponent responses to alternative actions taken from the same logged rollout state. During online learning and execution, both world models remain frozen and are queried by agents using only locally available information. Their prediction logits provide in-context guidance to an independent PPO policy. Across 30 matched seeds on Tag, Spread, and Adversary in the benchmark multi-particle environments, our proposed CASTLE achieves the highest mean final score among the evaluated methods, exceeding the strongest baseline on each task by 10.67, 6.46, and 0.33 normalized points, respectively.

Figures & tables

Explore similar work

CardsList
  1. Dreamer-CPC: Message Learning with World Models for Decentralized Multi-agent Reinforcement Learning

    Jul 22, 2026Taisuke Takayama, Naoto Yoshida, Tadahiro TaniguchiMulti-Agent Reinforcement LearningLatent States

  2. MA-JEPA: Joint-Embedding World Models for Multi-Agent Reinforcement Learning

    Sep 27, 2026Brandon Gary Kaplowitz, Osaze James Obahor, Christian Schroeder de WittMulti-Agent Reinforcement LearningWorld Models

  3. MAGIC: Multi-Step Advantage-Gated Causal Influence for Multi-agent Reinforcement Learning

    May 3, 2026Haohan Yu, Jinmiao Cong, Shengzhi Wang +2Multi-Agent Reinforcement LearningAction Expert