cs.CVOct 1, 2026

World Observer: Joint Actor-Observer Generation for Persistent World Modeling

Authors: Hyunwook Choi, Dahyun Chung, Hyunsung Kim, Siyoon Jin, Jinhyeok Choi, Junyoung Seo, Seungryong Kim

Organizations: KAIST AI

Abstract

How can a world model continuously observe regions beyond the actor's current view? Video world models simulate how an environment evolves from an agent's actions, yet remain actor-centric. Once an object leaves the actor's view, they lose direct evidence of its evolution, often failing to preserve its state and dynamics upon re-entry. To address this, we introduce World Observer, which decouples observing from acting by jointly generating a perspective actor for the agent-centric view with one or more panoramic observers that watch selected world regions. This allows objects that leave the actor's view to remain visually evolving in an observer, so their updated states are reflected when they re-enter. We ground the actor and observers by warping from a shared panoramic source for explicit geometric correspondence, and introduce an Observer Sink of high-resolution perspective references to restore fine appearance upon re-entry. Since the observers are decoupled from the actor, they can be placed freely across the scene, extended to multiple locations for broader coverage, and driven by control signals to steer out-of-view evolution. To evaluate out-of-view evolution, we further introduce world-space metrics and a benchmark spanning real and synthetic scenes. World Observer substantially improves out-of-view dynamics while remaining competitive in visual fidelity, camera control, and 3D adherence.

Figures & tables

Appendix figures & tables6 assets

Supplementary material from the paper’s appendix.

Appendix

Explore similar work

CardsList
  1. WorldDirector: Building Controllable World Simulators with Persistent Dynamic Memory

    Jul 2, 2026Hanlin Wang, Hao Ouyang, Qiuyu Wang +10Video World ModelsDramadirector

  2. Current World Models Lack a Persistent State Core

    Jun 18, 2026Jinpeng Lu, Dexu Zhu, Haoyuan Shi +8World ModelsLong-Horizon Inference

  3. World in World: Explore the World with World Models

    Sep 11, 2026Chenxi Song, Yanming Yang, Chi ZhangVideo World ModelsAutoregressive Video Generation