cs.CVOct 8, 2026

Multi-Agent Egocentric World Model with Fine-Grained Embodied Interaction

Authors: Dahyun Chung, Siyoon Jin, Hyunwook Choi, Honggyu An, Junyoung Seo, Hyunsung Kim, Seung Wook Kim, Seungryong Kim

Organizations: KAIST AI

Abstract

Egocentric world models predict first-person observations conditioned on an agent's actions, but most focus on a single agent. Real embodied settings often involve multiple agents that act and interact within a shared environment. Existing multi-agent world models rely on coarse actions like locomotion, camera control, or discrete commands, leaving fine-grained embodied interactions underexplored. We formulate multi-agent egocentric world modeling as synchronized ego-stream generation for multiple agents interacting through fine-grained actions in a shared world. This requires cross-view action consistency, shared-environment consistency, and consistent propagation of interaction-induced state updates. We propose Multi-agent Egocentric World Model (ME-World), which jointly denoises multiple ego streams in a shared token sequence, conditions each stream on all agents' target-view poses, and grounds generation with shared environment memory. We train and evaluate on real and synthetic multi-agent data and introduce shared-world consistency metrics for environment, update, and identity consistency. Experiments show ME-World improves shared-world consistency, action control, identity preservation, and video quality over existing methods.

Figures & tables

Appendix figures & tables7 assets

Supplementary material from the paper’s appendix.

Appendix

Explore similar work

CardsList
  1. MultiWorld: Scalable Multi-Agent Multi-View Video World Models

    Apr 20, 2026Haoyu Wu, Jiwen Yu, Yingtian Zou +1Multi-View ConsistencyAction-Conditioned World Models

  2. MetaWorld: Scaling Multi-Agent Video World Model from Single-view Video Data

    Jun 1, 2026Teng Hu, Mingchun Lu, Yating Wang +6Multi-View ConsistencyVideo Diffusion Models

  3. Gamma-World: Generative Multi-Agent World Modeling Beyond Two Players

    May 27, 2026Fangfu Liu, Kai He, Tianchang Shen +7Interactive World ModelsVideo World Models