MA-JEPA: Joint-Embedding World Models for Multi-Agent Reinforcement Learning
Organizations: Department of Engineering Science, University of Oxford
Abstract
World models improve sample efficiency by training policies on imagined trajectories, but their usefulness depends on learning representations that capture the information needed for future control. We study whether self-supervised joint-embedding prediction (JEPA) can provide this learning signal for multi-agent reinforcement learning. We introduce MA-JEPA, a stochastic world model that replaces observation reconstruction with prediction of target representations, enabling model-based multi-agent reinforcement learning with centralized training and decentralized execution. A categorical latent state and a causal Transformer are trained with posterior and action-conditioned dynamics prediction objectives and are then used for actor-critic learning from latent imagination. A training-only joint predictor conditions on all agents' local states and actions to predict each agent's next local observation embedding. These predictions are passed through the same local posterior used during real interaction with a centralized critic that is used only for value learning, with execution remaining decentralized. Our experiments show that this architecture performs strongly on SMAC, matching or exceeding the strongest reported comparator mean win rate on four of eight evaluated maps.
Figures & tables
| Map | Steps | MA-JEPA (ours) | DMAWM | MAPPO | QMIX | MAT |
| Easy 2s_vs_1sc | 100k | 97.3 (2.5) | 95.7 (5.9) | 100.0 (0.0) | 13.7 (7.4) | 51.3 (27.1) |
| 2s3z | 100k | 78.7 (14.5) | 88.0 (7.0) | 8.0 (2.6) | 31.3 (24.5) | 16.5 (9.2) |
| 8m | 100k | 92.0 (1.0) | 92.0 (4.6) | 65.7 (33.1) | 83.0 (10.1) | 26.0 (22.6) |
| MMM | 100k | 91.0 (3.0) | 90.0 (4.6) | 20.7 (17.0) | 3.0 (2.6) | 0.0 (0.0) |
| so_many_baneling | 100k | 95.0 (6.1) | 99.5 (0.7) | 77.0 (7.2) | 21.7 (18.0) | 32.0 (24.3) |
| Hard 3s_vs_4z | 100k | 74.0 (7.5) | 95.7 (2.3) | 0.0 (0.0) | 0.0 (0.0) | 0.0 (0.0) |
Appendix figures & tables8 assets
Supplementary material from the paper’s appendix.
Appendix
| Aspect | MARIE | MAMBA | MBVD | MATWM | DMAWM | MA-JEPA |
| Multi-Agent Design | Yes | Yes | Yes | Yes | Yes | Yes |
| Backbone Architecture | Transformer | GRU | GRU | Transformer | GRU + Transformer | Transformer |
| Latent Representation | VQ-VAE | Categorical VAE | VAE | Categorical VAE | Categorical RSSM | Categorical latent |
| Critic Type | Centralized | Centralized | Centralized | Semi-centralized | Centralized | Centralized |
| Agent Training | PPO-style | PPO-style | Deep Q-learning | DreamerV3-style | PPO-style | PPO-style |
| Scenario | Difficulty | Training budget |
| 2m_vs_1z | Easy | 100K |
| 2s_vs_1sc | Easy | 100K |
| 2s3z | Easy | 100K |
| 3m | Easy | 100K |
| 3s_vs_3z | Easy | 100K |
| 8m | Easy | 100K |
| Hyperparameter | Value used |
| World model | |
| Observation encoder layers / width | 3 / 1024 |
| Deterministic state dimension | 4096 |
| Categorical latent | |
| Latent uniform mixture | 0.01 |
| Local Transformer layers / width / heads | 2 / 512 / 8 |
| Hyperparameter | Value used |
| Reinforcement learning | |
| Optimizer | Adam |
| Actor / critic learning rate | / |
| Actor / critic gradient clipping | 100 / 100 |
| Entropy coefficient | 0.01 |
| PPO epochs / clipping | 5 / 0.2 |
| Hyperparameter | MAPPO | MAT |
| Policy | Recurrent or FF | Transformer |
| Optimizer | Adam | Adam |
| Actor / critic learning rate | each | Joint |
| Rollout length per collector | 400 | 100 |
| Parallel collection environments | 8 | 32 |
| Transitions per rollout batch | 3200 | 3200 |
| Hyperparameter | Value used |
| Agent architecture / hidden width | Recurrent / 64 |
| Optimizer / learning rate | Adam / |
| Discount / TD | 0.99 / 0.6 |
| Target-network update interval | 200 episodes |
| Maximum gradient norm | 10 |
| Training batch size | 128 episodes |
| Hyperparameter | Value used |
| World-model / actor / critic optimizer | Adam / Adam / Adam |
| World-model learning rate | |
| Actor / critic learning rate | / |
| Actor weight decay | |
| World-model / policy gradient clipping | 100 / 100 |
| Replay capacity / minimum replay | / 500 |