cs.CVMar 20, 2026

EgoForge: Goal-Directed Egocentric World Simulator

Authors: Yifan Shen, Jiateng Liu, Xinzhuo Li, Yuanzhe Liu, Bingxuan Li, Houze Yang, Wenqi Jia, Yijiang Li, +4 more

Organizations: University of Illinois Urbana-Champaign · University of California San Diego

Abstract

Generative world models have shown promise for simulating dynamic environments, yet egocentric video remains challenging due to rapid viewpoint changes, frequent hand-object interactions, and goal-directed procedures whose evolution depends on latent human intent. Existing approaches either focus on hand-centric instructional synthesis with limited scene evolution, perform static view translation without modeling action dynamics, or rely on dense supervision, such as camera trajectories, long video prefixes, and synchronized multi-camera capture. In this work, we introduce EgoForge, an egocentric goal-directed world simulator that generates coherent, first-person video rollouts from minimal static inputs: a single egocentric image, a high-level instruction, and an optional auxiliary exocentric view. To improve intent alignment and temporal coherence, we introduce GRAFT, a trajectory-level diffusion refinement method that uses positive and negative rollout distributions, derived from goal, temporal, scene-consistency, and perceptual rewards, to steer the diffusion velocity field toward coherent, goal-complete egocentric simulations. Extensive experiments show EgoForge achieves consistent gains in semantic alignment, geometric stability, and motion fidelity over strong baselines, and performs robustly in real-world smart-glasses experiments.

Figures & tables

Appendix figures & tables5 assets

Supplementary material from the paper’s appendix.

Appendix

Explore similar work

CardsList
  1. EgoSim: Egocentric World Simulator for Embodied Interaction Generation

    Apr 1, 2026Jinkun Hao, Mingda Jia, Ruiyan Wang +7Cross-Embodiment TransferEmbodied

  2. EgoGenesis: Egocentric World-Action Modeling with Online Anchored Projective Memory and Action-3D RoPE

    Jul 30, 2026Zexuan Yan, Yuzhou Wu, Yue Ma +9Egocentric VideoCamera-Conditioned Video Generation

  3. Ego-Forge: Text and Geometric-Attention Free Exo-to-Egocentric Video Generation

    Sep 28, 2026Mohammad Mahdi, Luc Van Gool, Danda Pani PaudelEgocentric VideoEgocentric Vision