Egocentric world models predict first-person observations conditioned on an agent's actions, but most focus on a single agent. Real embodied settings often involve multiple agents that act and interact within a shared environment. Existing multi-agent world models rely on coarse actions like locomotion, camera control, or discrete commands, leaving fine-grained embodied interactions underexplored. We formulate multi-agent egocentric world modeling as synchronized ego-stream generation for multiple agents interacting through fine-grained actions in a shared world. This requires cross-view action consistency, shared-environment consistency, and consistent propagation of interaction-induced state updates. We propose Multi-agent Egocentric World Model (ME-World), which jointly denoises multiple ego streams in a shared token sequence, conditions each stream on all agents' target-view poses, and grounds generation with shared environment memory. We train and evaluate on real and synthetic multi-agent data and introduce shared-world consistency metrics for environment, update, and identity consistency. Experiments show ME-World improves shared-world consistency, action control, identity preservation, and video quality over existing methods.
Figures & tables
Figure 1: Teaser. Given multi-agent motions, head poses, character identities, and a shared world with first-frame initialization in (a), ME-World jointly generates synchronized ego streams (b) that remain consistent in character identity, scene appearance, and interaction outcomes across agents.
Figure 2: Motivation. Three challenges in multi-agent egocentric world modeling. Without shared action conditioning, (a) actions are inconsistent across ego streams. Without joint generation, independent ego streams show (b) environment inconsistency and (c) inconsistent state updates.
Figure 3: Overall architecture (a) For each agent, ME-World builds viewing ray conditions and shared-action conditions from agents’ input and build shared environment memory from the both agents’ history. (b) ME-World jointly generates synchronized ego streams.
Figure 4: Evaluation Protocol for Shared-World Consistency. We evaluate generated environments with Senv , state updates with Supdate , and identity with Sid . Ground-truth geometry is used only to establish co-visible correspondences, while consistency is measured on generated streams.
Figure 5: Qualitative Comparison with Existing Methods. We compare ME-World on the real (top) and synthetic (bottom) benchmarks. Each column shows three timestamps from two synchronized ego streams. For multi-view baselines, the ground-truth conditioning stream is shown with reduced opacity. Best viewed zoomed in.
Figure 6: Qualitative Comparison with Multi-Agent World Models. We compare ME-World with existing multi-agent world models on the real benchmark. Each column shows three timestamps from two synchronized ego streams. ME-World better preserves a coherent shared environment and synchronized interactions across the two viewpoints. Best viewed zoomed in.
Shared-World Consistency
Camera Control
Self Action Control
Other-Agent Action Control
Video Quality
Method
Senv↑
Supdate↑
Sid↑
Trans. Err. ↓
Rot. Err. ↓
Hand-F1 ↑
Hand-mIoU ↑
Full-Body-F1 ↑
Full-Body-PCK ↑
PSNR ↑
SSIM ↑
LPIPS ↓
FVD ↓
Real Benchmark
GEN3C ( Ren et al., 2025 )
0.417
0.412
N/A
0.12
12.376
0.537
0.087
0.229
0.456
18.083
0.582
0.444
1358
AnyView-DVS ( Van Hoorick et al., 2026 )
0.332
0.319
N/A
0.079
4.736
0.169
0.005
0.037
0.071
15.899
0.546
0.603
1834.9
EgoSim ( Hao et al., 2026 )
0.332
0.319
N/A
0.114
11.501
0.75
0.393
0.274
0.514
16.398
0.574
0.448
1012.7
JointControlVideo ( Zhang et al., 2026 )
0.328
0.291
N/A
0.166
21.663
0.799
0.249
0.215
0.402
15.18
0.473
0.574
694
Table 1: Quantitative Comparison with Existing Methods. We compare ME-World with multi-view generation, single-ego world, and general world models on the real and synthetic benchmarks. N/A denotes metrics that are not applicable to a benchmark.
Shared-World Consistency
Camera Control
Self Action Control
Other-Agent Action Control
Video Quality
Method
Senv↑
Supdate↑
Trans. Err. ↓
Rot. Err. ↓
Hand-F1 ↑
Hand-mIoU ↑
Full-Body-F1 ↑
Full-Body-PCK ↑
PSNR ↑
SSIM ↑
LPIPS ↓
FVD ↓
Real Benchmark
MultiWorld ( Wu et al., 2026 )
0.395
0.376
0.113
11.616
0.868
0.497
0.804
0.837
15.323
0.481
0.525
694
Solaris ( Savva et al., 2026 )
0.324
0.293
0.171
27.18
0.709
0.048
0.146
0.236
14.98
0.474
0.594
984.2
γ -World ( Liu et al., 2026 )
0.304
0.266
0.181
24.414
0.758
0.064
0.188
0.214
14.892
0.469
0.584
788
MetaWorld ( Hu et al., 2026b )
0.421
0.413
0.054
4.086
0.865
0.379
0.816
0.787
18.928
0.627
0.329
704.4
Table 2: Quantitative Comparison with Multi-Agent World Models. We compare ME-World with multi-agent world models on the real benchmark. Since existing multi-agent world models are designed for different domains and control interfaces, we adapt them to our embodied multi-agent setting for a controlled comparison.
Method
Joint Multi-Agent Generation
Shared Action Conditioning
Shared Environment Memory
MultiWorld ( Wu et al., 2026 )
✓
✓
△
Solaris ( Savva et al., 2026 )
✓
✓
×
γ -World ( Liu et al., 2026 )
△
✓
×
MetaWorld ( Hu et al., 2026b )
△
✓
△
ME-World
✓
✓
✓
Table 3: Adapted Multi-Agent Architectures. ✓ uses the corresponding ME-World component, △ denotes a method-specific replacement, and × no explicit counterpart.
Shared-World Consistency
Camera Control
Self Action Control
Other-Agent Action Control
Video Quality
Index
Variant
FT
Senv↑
Supdate↑
Trans. Err. ↓
Rot. Err. ↓
Hand-F1 ↑
Hand-mIoU ↑
Full-Body-F1 ↑
Full-Body-PCK ↑
PSNR ↑
SSIM ↑
LPIPS ↓
FVD ↓
Real Benchmark
(1)
Cosmos-Predict2.5 (I2V + Text)
×
0.289
0.245
0.178
19.993
0.783
0.080
0.248
0.284
14.701
0.468
0.586
1006
(2)
Cosmos-Predict2.5 (I2V + Text)
✓
0.290
0.249
0.170
24.513
0.789
0.079
0.238
0.309
14.941
0.473
0.602
768.3
(3)
(2) + Self Hand Pose + History Warp
✓
0.401
0.388
0.056
4.998
0.857
0.457
0.580
0.385
18.835
0.628
0.349
558.6
(4)
(3) + Joint Gen.
✓
0.426
0.421
0.067
6.451
0.846
0.354
0.563
0.416
18.985
0.637
0.344
554.7
Table 4: Component Ablation Results. We evaluate the effects of each component. ’FT’ means fine-tuned by our real dataset.
Figure 7: Qualitative Results of Ablation Studies. Removing individual components leads to inconsistencies in agent appearance, embodied actions, and the shared environment across ego streams. Red boxes highlight representative artifacts. Best viewed zoomed in.
Appendix figures & tables7 assets
Supplementary material from the paper’s appendix.
Appendix
Figure 8: Additional Qualitative Results on the Real Benchmark. We show additional generations from diverse interaction sequences in real benchmark. Each example contains the two synchronized ego streams generated jointly by our model across multiple timesteps. Best viewed zoomed in.
Figure 9: Additional Qualitative Results on the Synthetic Benchmark. We show additional generations from diverse interaction sequences in synthetic benchmark. For each example, the two ego views are jointly generated by our model. The third-person view (TPV) is shown only as a visual reference to help interpret the shared scene and interaction, and is not generated or used as input by our model. Best viewed zoomed in.
Figure 10: Multi-Agent and Autoregressive Generation. We further demonstrate our model in more diverse generation settings, including multi-agent and long-horizon generation. Our model jointly generates three synchronized ego views for three interacting agents (top), and autoregressively generates two-agent sequences of 221 frames (bottom). For the three-agent results, the third-person view (TPV) is provided only as a visual reference to help interpret the shared scene and interaction, and is not generated or used as input by our model. Best viewed zoomed in.
Figure 11: Metric validation. We visualize the query regions evaluated by Senv and Supdate , their co-visible correspondences in the target stream, and the resulting DINOv3 cosine similarities. The consistent generation in (a) shows high correspondence similarity and receives higher metric scores, whereas the cross-view inconsistencies in (b) appear as localized similarity drops and lower scores. This correspondence provides qualitative evidence that the proposed metrics capture the intended cross-view consistency. Best viewed zoomed in.
Metric
Ours
Misalignment
Env. Conflict
State Conflict
Senv
0.468
0.399 ( −0.069 )
0.377 ( −0.090 )
0.408 ( −0.059 )
Supdate
0.466
0.377 ( −0.090 )
0.447 ( −0.019 )
0.320 ( −0.147 )
Appendix
Table 5: Controlled validation of the proposed metrics. We evaluate how Senv and Supdate respond to temporal misalignment, environment conflict, and state conflict. Values in brackets denote changes from the original generations.
Figure 12: Dataset example. Examples are from (a) our generated synthetic dataset and (b) curated real-world CoMind dataset, showing synchronized ego streams and corresponding human poses.
Figure 13: Synthetic dataset generation pipeline. (a) We collect diverse 3D characters, environments, and paired motions, (b) apply the motions to characters, place them in the shared world space, and attach egocentric cameras, and (c) render synchronized egocentric video pairs.