Egocentric world models predict first-person observations conditioned on an agent's actions, but most focus on a single agent. Real embodied settings often involve multiple agents that act and interact within a shared environment. Existing multi-agent world models rely on coarse actions like locomotion, camera control, or discrete commands, leaving fine-grained embodied interactions underexplored. We formulate multi-agent egocentric world modeling as synchronized ego-stream generation for multiple agents interacting through fine-grained actions in a shared world. This requires cross-view action consistency, shared-environment consistency, and consistent propagation of interaction-induced state updates. We propose Multi-agent Egocentric World Model (ME-World), which jointly denoises multiple ego streams in a shared token sequence, conditions each stream on all agents' target-view poses, and grounds generation with shared environment memory. We train and evaluate on real and synthetic multi-agent data and introduce shared-world consistency metrics for environment, update, and identity consistency. Experiments show ME-World improves shared-world consistency, action control, identity preservation, and video quality over existing methods.
Figures & tables
Figure 1: Teaser. Given multi-agent motions, head poses, character identities, and a shared world with first-frame initialization in (a), ME-World jointly generates synchronized ego streams (b) that remain consistent in character identity, scene appearance, and interaction outcomes across agents.
Figure 2: Motivation. Three challenges in multi-agent egocentric world modeling. Without shared action conditioning, (a) actions are inconsistent across ego streams. Without joint generation, independent ego streams show (b) environment inconsistency and (c) inconsistent state updates.
Figure 3: Overall architecture (a) For each agent, ME-World builds viewing ray conditions and shared-action conditions from agents’ input and build shared environment memory from the both agents’ history. (b) ME-World jointly generates synchronized ego streams.
Figure 4: Evaluation Protocol for Shared-World Consistency. We evaluate generated environments with Senv , state updates with Supdate , and identity with Sid . Ground-truth geometry is used only to establish co-visible correspondences, while consistency is measured on generated streams.
Figure 5: Qualitative Comparison with Existing Methods. We compare ME-World on the real (top) and synthetic (bottom) benchmarks. Each column shows three timestamps from two synchronized ego streams. For multi-view baselines, the ground-truth conditioning stream is shown with reduced opacity. Best viewed zoomed in.
Figure 6: Qualitative Comparison with Multi-Agent World Models. We compare ME-World with existing multi-agent world models on the real benchmark. Each column shows three timestamps from two synchronized ego streams. ME-World better preserves a coherent shared environment and synchronized interactions across the two viewpoints. Best viewed zoomed in.
Shared-World Consistency
Camera Control
Self Action Control
Other-Agent Action Control
Video Quality
Method
Senv↑
Supdate↑
Sid↑
Trans. Err. ↓
Rot. Err. ↓
Hand-F1 ↑
Hand-mIoU ↑
Full-Body-F1 ↑
Full-Body-PCK ↑
PSNR ↑
SSIM ↑
LPIPS ↓
FVD ↓
Real Benchmark
GEN3C ( Ren et al., 2025 )
0.417
0.412
N/A
0.12
12.376
0.537
0.087
0.229
0.456
18.083
0.582
0.444
1358
AnyView-DVS ( Van Hoorick et al., 2026 )
0.332
0.319
N/A
0.079
4.736
0.169
0.005
0.037
0.071
15.899
0.546
0.603
1834.9
EgoSim ( Hao et al., 2026 )
0.332
0.319
N/A
0.114
11.501
0.75
0.393
0.274
0.514
16.398
0.574
0.448
1012.7
JointControlVideo ( Zhang et al., 2026 )
0.328
0.291
N/A
0.166
21.663
0.799
0.249
0.215
0.402
15.18
0.473
0.574
694
Table 1: Quantitative Comparison with Existing Methods. We compare ME-World with multi-view generation, single-ego world, and general world models on the real and synthetic benchmarks. N/A denotes metrics that are not applicable to a benchmark.
Shared-World Consistency
Camera Control
Self Action Control
Other-Agent Action Control
Video Quality
Method
Senv↑
Supdate↑
Trans. Err. ↓
Rot. Err. ↓
Hand-F1 ↑
Hand-mIoU ↑
Full-Body-F1 ↑
Full-Body-PCK ↑
PSNR ↑
SSIM ↑
LPIPS ↓
FVD ↓
Real Benchmark
MultiWorld ( Wu et al., 2026 )
0.395
0.376
0.113
11.616
0.868
0.497
0.804
0.837
15.323
0.481
0.525
694
Solaris ( Savva et al., 2026 )
0.324
0.293
0.171
27.18
0.709
0.048
0.146
0.236
14.98
0.474
0.594
984.2
γ -World ( Liu et al., 2026 )
0.304
0.266
0.181
24.414
0.758
0.064
0.188
0.214
14.892
0.469
0.584
788
MetaWorld ( Hu et al., 2026b )
0.421
0.413
0.054
4.086
0.865
0.379
0.816
0.787
18.928
0.627
0.329
704.4
Table 2: Quantitative Comparison with Multi-Agent World Models. We compare ME-World with multi-agent world models on the real benchmark. Since existing multi-agent world models are designed for different domains and control interfaces, we adapt them to our embodied multi-agent setting for a controlled comparison.
Method
Joint Multi-Agent Generation
Shared Action Conditioning
Shared Environment Memory
MultiWorld ( Wu et al., 2026 )
✓
✓
△
Solaris ( Savva et al., 2026 )
✓
✓
×
γ -World ( Liu et al., 2026 )
△
✓
×
MetaWorld ( Hu et al., 2026b )
△
✓
△
ME-World
✓
✓
✓
Table 3: Adapted Multi-Agent Architectures. ✓ uses the corresponding ME-World component, △ denotes a method-specific replacement, and × no explicit counterpart.
Shared-World Consistency
Camera Control
Self Action Control
Other-Agent Action Control
Video Quality
Index
Variant
FT
Senv↑
Supdate↑
Trans. Err. ↓
Rot. Err. ↓
Hand-F1 ↑
Hand-mIoU ↑
Full-Body-F1 ↑
Full-Body-PCK ↑
PSNR ↑
SSIM ↑
LPIPS ↓
FVD ↓
Real Benchmark
(1)
Cosmos-Predict2.5 (I2V + Text)
×
0.289
0.245
0.178
19.993
0.783
0.080
0.248
0.284
14.701
0.468
0.586
1006
(2)
Cosmos-Predict2.5 (I2V + Text)
✓
0.290
0.249
0.170
24.513
0.789
0.079
0.238
0.309
14.941
0.473
0.602
768.3
(3)
(2) + Self Hand Pose + History Warp
✓
0.401
0.388
0.056
4.998
0.857
0.457
0.580
0.385
18.835
0.628
0.349
558.6
(4)
(3) + Joint Gen.
✓
0.426
0.421
0.067
6.451
0.846
0.354
0.563
0.416
18.985
0.637
0.344
554.7
Table 4: Component Ablation Results. We evaluate the effects of each component. ’FT’ means fine-tuned by our real dataset.
Figure 7: Qualitative Results of Ablation Studies. Removing individual components leads to inconsistencies in agent appearance, embodied actions, and the shared environment across ego streams. Red boxes highlight representative artifacts. Best viewed zoomed in.
Appendix figures & tables7 assets
Supplementary material from the paper’s appendix.
Appendix
Figure 8: Additional Qualitative Results on the Real Benchmark. We show additional generations from diverse interaction sequences in real benchmark. Each example contains the two synchronized ego streams generated jointly by our model across multiple timesteps. Best viewed zoomed in.
Figure 9: Additional Qualitative Results on the Synthetic Benchmark. We show additional generations from diverse interaction sequences in synthetic benchmark. For each example, the two ego views are jointly generated by our model. The third-person view (TPV) is shown only as a visual reference to help interpret the shared scene and interaction, and is not generated or used as input by our model. Best viewed zoomed in.
Figure 10: Multi-Agent and Autoregressive Generation. We further demonstrate our model in more diverse generation settings, including multi-agent and long-horizon generation. Our model jointly generates three synchronized ego views for three interacting agents (top), and autoregressively generates two-agent sequences of 221 frames (bottom). For the three-agent results, the third-person view (TPV) is provided only as a visual reference to help interpret the shared scene and interaction, and is not generated or used as input by our model. Best viewed zoomed in.
Figure 11: Metric validation. We visualize the query regions evaluated by Senv and Supdate , their co-visible correspondences in the target stream, and the resulting DINOv3 cosine similarities. The consistent generation in (a) shows high correspondence similarity and receives higher metric scores, whereas the cross-view inconsistencies in (b) appear as localized similarity drops and lower scores. This correspondence provides qualitative evidence that the proposed metrics capture the intended cross-view consistency. Best viewed zoomed in.
Metric
Ours
Misalignment
Env. Conflict
State Conflict
Senv
0.468
0.399 ( −0.069 )
0.377 ( −0.090 )
0.408 ( −0.059 )
Supdate
0.466
0.377 ( −0.090 )
0.447 ( −0.019 )
0.320 ( −0.147 )
Appendix
Table 5: Controlled validation of the proposed metrics. We evaluate how Senv and Supdate respond to temporal misalignment, environment conflict, and state conflict. Values in brackets denote changes from the original generations.
Figure 12: Dataset example. Examples are from (a) our generated synthetic dataset and (b) curated real-world CoMind dataset, showing synchronized ego streams and corresponding human poses.
Figure 13: Synthetic dataset generation pipeline. (a) We collect diverse 3D characters, environments, and paired motions, (b) apply the motions to characters, place them in the shared world space, and attach egocentric cameras, and (c) render synchronized egocentric video pairs.
Video world models have achieved remarkable success in simulating environmental dynamics in response to actions by users or agents. They are modeled as action-conditioned video generation models that take historical frames and current actions as input to predict future frames. Yet, most existing approaches are limited to single-agent scenarios and fail to capture the complex interactions inherent in real-world multi-agent systems. We present \textbf{MultiWorld}, a unified framework for multi-agent multi-view world modeling that enables accurate control of multiple agents while maintaining multi-view consistency. We introduce the Multi-Agent Condition Module to achieve precise multi-agent controllability, and the Global State Encoder to ensure coherent observations across different views. MultiWorld supports flexible scaling of agent and view counts, and synthesizes different views in parallel for high efficiency. Experiments on multi-player game environments and multi-robot manipulation tasks demonstrate that MultiWorld outperforms baselines in video fidelity, action-following ability, and multi-view consistency. Project page: https://multi-world.github.io/
Video world models are a foundational generative technology for embodied AI and the Metaverse, yet existing approaches are inherently limited to a single agent observing from a single perspective. Extending these models to multi-agent settings introduces two critical challenges: data scarcity (coordinated multi-view recordings are prohibitively expensive to collect for general open-domain scenarios) and world state alignment (independently generated video streams cannot ensure that shared physical environments and events evolve consistently across views). To address these challenges, we propose MetaWorld, a novel framework that scales multi-agent video world models to open-domain environments directly from single-view videos. First, we introduce Monocular World-State Unrolling (MWSU) to explicitly decompose monocular footage into the camera operator's ego-motion and the visible subject's spatial trajectory. This camera-trajectory decomposition naturally extracts synchronized multi-agent motion data within a shared 3D space, completely bypassing the need for multi-camera setups. Second, for precise visual control, we develop the Subject-Aware World Generator to enable appearance-driven simulation conditioned on per-agent identity images. Finally, to ensure both views are grounded in the identical physical reality, we propose World-State Alignment, a per-frame inter-branch cross-attention mechanism inserted at every transformer layer of the video DiT. By jointly synchronizing the denoising process, WSA enforces both static geometric consistency and dynamic motion consistency, encouraging that the shared 3D environment and physical events remain well-aligned across both egocentric views. Extensive experiments demonstrate that MetaWorld achieves superior cross-view consistency and identity fidelity, establishing a highly scalable, physics-driven paradigm for multi-agent video world modeling.
Teng Hu, Mingchun Lu, Yating Wang +6
Shanghai Jiao Tong University · Zhejiang University · Nanyang Technological University
World models for interactive video generation have largely focused on single-agent settings, where future observations are generated from a single control signal. However, many generated environments require multi-agent interaction: multiple players, robots, or embodied agents act simultaneously within a shared space. Scaling world models to such settings requires a principled multi-agent design: agents should remain independently controllable, permutation-symmetric, and support efficient inference while maintaining consistency across time and perspectives. In this paper, we present our generative multi-agent world model for interactive simulation. It introduces Simplex Rotary Agent Encoding, a parameter-free extension of 3D RoPE that represents agents as vertices of a regular simplex in rotary angle space. This gives each agent a distinct phase while making all agents permutation-equivalent, enabling scalable agent identity without learned per-slot identities or a fixed agent ordering. To avoid dense all-to-all attention across agents, we further propose Sparse Hub Attention, where learnable hub tokens mediate token interaction across agents, reducing cross-agent attention cost from quadratic to linear in the number of agents. For real-time rollout, we distill a full-context diffusion teacher into a causal student that generates temporal blocks sequentially with KV caching, enabling action-responsive generation at 24 FPS. Experiments in multiplayer virtual environments show that our model improves video fidelity, action controllability, and inter-agent consistency over slot-based and dense-attention baselines, while generalizing from two to four players without additional training.
Fangfu Liu, Kai He, Tianchang Shen +7
NVIDIA · Tsinghua University · University of Toronto +1