Generative world models have shown promise for simulating dynamic environments, yet egocentric video remains challenging due to rapid viewpoint changes, frequent hand-object interactions, and goal-directed procedures whose evolution depends on latent human intent. Existing approaches either focus on hand-centric instructional synthesis with limited scene evolution, perform static view translation without modeling action dynamics, or rely on dense supervision, such as camera trajectories, long video prefixes, and synchronized multi-camera capture. In this work, we introduce EgoForge, an egocentric goal-directed world simulator that generates coherent, first-person video rollouts from minimal static inputs: a single egocentric image, a high-level instruction, and an optional auxiliary exocentric view. To improve intent alignment and temporal coherence, we introduce GRAFT, a trajectory-level diffusion refinement method that uses positive and negative rollout distributions, derived from goal, temporal, scene-consistency, and perceptual rewards, to steer the diffusion velocity field toward coherent, goal-complete egocentric simulations. Extensive experiments show EgoForge achieves consistent gains in semantic alignment, geometric stability, and motion fidelity over strong baselines, and performs robustly in real-world smart-glasses experiments.
Figures & tables
Figure 1 : EgoForge Overview : Given a single egocentric observation, a high-level instruction expressing user intent, and an auxiliary exo-view reference, EgoForge fuses encoded visual features with noisy video latents at each DiT block to guide generation. Geometry alignment weakly supervises intermediate features using angular and scale consistency to encourage spatially stable rollouts. The resulting rollout videos are further refined via the proposed GRAFT alignment that optimizes goal completion, scene consistency, temporal causality, and perceptual fidelity.
Figure 2 : GRAFT refinement. Rollout-level rewards yield normalized weights r and 1−r for the implicit positive and negative branch losses. Only the trainable model receives gradients, while the old model is refreshed across iterations.
Model
DINO-Score ↑
CLIP-Score ↑
SSIM ↑
LPIPS ↓
FVD ↓
flow MSE ↓
PSNR ↑
EgoDreamer [ 61 ]
42.35
25.40
0.58
0.35
580.45
8.15
15.20
Handi [ 31 ]
31.12
18.25
0.42
0.52
912.30
14.50
12.85
Cosmos [ 42 ]
49.42
29.77
0.70
0.26
448.12
6.40
18.73
HunyuanVideo [ 28 ]
53.54
29.43
0.71
0.26
384.31
6.10
18.88
WAN2.2 [ 58 ]
53.99
35.69
0.72
0.23
322.17
5.78
20.44
EgoForge (Ours)
61.25
39.30
0.79
0.15
182.25
2.83
24.08
Table 2 : Quantitative comparisons on the X-Ego benchmark. EgoForge outperforms all baselines across semantic, perceptual, and temporal metrics.
Model
DINO-Score ↑
CLIP-Score ↑
SSIM ↑
LPIPS ↓
FVD ↓
flow MSE ↓
PSNR ↑
Cosmos+EV
48.60
29.60
0.67
0.28
485.75
6.82
18.30
Cosmos+TT
50.80
30.40
0.71
0.25
433.90
6.31
18.88
HunyuanVideo+EV
52.80
29.20
0.70
0.27
405.87
6.30
18.61
HunyuanVideo+TT
54.10
29.86
0.72
0.24
365.80
5.95
19.10
WAN2.2+EV
52.91
35.11
0.71
0.27
352.41
6.25
20.05
WAN2.2+TT
54.80
36.20
0.73
0.25
310.57
5.60
20.64
Table 3 : Quantitative comparisons on X-Ego between EgoForge and other finetuned baseline variants. +EV: with exo-view img, +TT: text-only domain adaptation; +CI: our conditioning inputs (exo-view and goal instructions) with our structured injection with Geometry Weak Supervision.
FT
GWS
GRAFT
DINO-Score ↑
CLIP-Score ↑
SSIM ↑
LPIPS ↓
FVD ↓
flow MSE ↓
PSNR ↑
✓
✗
✗
56.81
37.10
0.74
0.21
260.89
4.82
21.92
✓
✓
✗
58.92
38.05
0.76
0.18
218.72
3.92
22.87
✓
✓
✓
61.25
39.30
0.79
0.15
182.25
2.83
24.08
Table 4 : Ablation on EgoForge modules. We evaluate the impact of diffusion finetuning (FT), geometry weak supervision (GWS), and the proposed GRAFT refinement. Each component consistently improves performance across semantic, perceptual, and temporal metrics, with the full EgoForge model achieving the best overall results.
Conditioning
Inputs
Generation Quality
Ego
Text
Exo
DINO ↑
CLIP ↑
SSIM ↑
LPIPS ↓
FVD ↓
Flow MSE ↓
PSNR ↑
Single-modality conditioning
Text only
✗
✓
✗
49.84
36.75
0.62
0.33
398.52
7.01
17.18
Ego only
✓
✗
✗
54.26
28.41
0.70
0.25
310.44
5.73
20.14
Exo only
✗
✗
✓
50.91
27.86
0.65
0.30
365.73
6.38
18.56
Two-modality conditioning
Table 5 : Ablation of conditioning inputs. Missing modalities are replaced with learned null embeddings. For counterfactual conditions, random Exo uses an exocentric view sampled from another test example, and wrong Text uses an action-incompatible instruction. Metrics computed against original held-out video and instruction.
Figure 3 : Qualitative Comparison between EgoForge and baselines . EgoForge accurately reconstructs multi-step, causally ordered actions, preserving hand–object geometry, temporal consistency, and goal alignment. For instance, in the first example, Cosmos erroneously generates a third hand, Hunyuan depicts a disconnected arm, and Wan2.2 fails to complete the coffee-pouring task. In the second example, Cosmos generates multiple balls, while both Hunyuan and Wan2.2 generate an incorrect person to perform the action. In contrast, EgoForge accurately completes both tasks.
Figure 4 : Qualitative comparison with and without exocentric input. EgoForge can be steered using an auxiliary exo-view to improve spatial grounding by anchoring the simulation to the reference environment.
Appendix figures & tables5 assets
Supplementary material from the paper’s appendix.
Appendix
Rewards
DINO-Score ↑
CLIP-Score ↑
SSIM ↑
LPIPS ↓
FVD ↓
flow MSE ↓
PSNR ↑
✗ \penalty\penaltyRgoal
59.62
38.49
0.78
0.16
205.96
3.48
23.48
✗ \penalty\penaltyRenv
60.67
39.05
0.78
0.16
200.49
3.43
23.60
✗ \penalty\penaltyRtemp
60.78
39.11
0.78
0.16
213.25
3.70
23.72
✗ \penalty\penaltyRper
60.32
38.80
0.77
0.18
204.13
3.48
23.17
EgoForge (Ours)
61.25
39.30
0.79
0.15
182.25
2.83
24.08
Appendix
Table 16 : Effect of reward components in GRAFT . Each reward term contributes to improved semantic alignment and temporal coherence, while the full reward composition yields the strongest performance across all metrics.
Figure 5 : Qualitative egocentric video rollouts. EgoForge generates temporally coherent first-person video trajectories that follow the intended activity while preserving scene structure and realistic hand–object interactions across diverse environments.
Figure 6 : Qualitative Comparison between EgoForge and baselines . Sample frames from generated videos illustrating two scenarios. Top row: In the hand-washing task, baselines struggle with object consistency ( e.g. , Cosmos hallucinates the soap source) or ignore scene context ( e.g. , Wan2.2 bbypasses the soap on the table), while EgoForge successfully executes the action. using existing objects. Bottom row: In the soccer task, baselines exhibit severe artifacts like ghosting (Cosmos) or fail to follow precise instructions regarding motion and goals (Hunyuan, Wan2.2). EgoForge accurately executes the complex command: trapping with the left leg and shooting with the right.
Figure 7 : Dense visualization of long-duration sequences. For each EgoForge-generated video, we show 26 uniformly sampled frames to illustrate how actions unfold over time. Across diverse first-person tasks, EgoForge preserves scene structure, maintains a stable egocentric viewpoint, and produces smooth, goal-consistent hand–object interactions.
We introduce EgoSim, a closed-loop egocentric world simulator that generates spatially consistent interaction videos and persistently updates the underlying 3D scene state for continuous simulation. Existing egocentric simulators either lack explicit 3D grounding, causing structural drift under viewpoint changes, or treat the scene as static, failing to update world states across multi-stage interactions. EgoSim addresses both limitations by modeling 3D scenes as updatable world states. We generate embodiment interactions via a Geometry-action-aware Observation Simulation model, with spatial consistency from an Interaction-aware State Updating module. To overcome the critical data bottleneck posed by the difficulty in acquiring densely aligned scene-interaction training pairs, we design a scalable pipeline that extracts static point clouds, camera trajectories, and embodiment actions from in-the-wild large-scale monocular egocentric videos. We further introduce EgoCap, a capture system that enables low-cost real-world data collection with uncalibrated smartphones. Extensive experiments demonstrate that EgoSim significantly outperforms existing methods in terms of visual quality, spatial consistency, and generalization to complex scenes and in-the-wild dexterous interactions, while supporting cross-embodiment transfer to robotic manipulation. Codes and datasets will be open soon. The project page is at egosimulator.github.io.
Jinkun Hao, Mingda Jia, Ruiyan Wang +7
Shanghai Jiaotong University · Shanghai AI Laboratory · The University of Hong Kong
Egocentric video offers rich manipulation experience for embodied AI, yet collecting diverse egocentric data across scenes, objects, motions, and embodiments remains costly. We present \method, an egocentric world-action simulator that synthesizes controllable, high-quality manipulation videos to expand scarce real-world training data. \method{} builds on a pretrained video generation prior and introduces two geometry-aware conditioning mechanisms. Online Anchored Projective Memory (OAPM) preserves a first-frame 3D scene anchor while periodically refreshing a recent state during autoregressive generation. Action-3D Rotary Position Embedding (A3D-RoPE) encodes end-effector motion with camera-aware 3D rotary coordinates, injecting action geometry into skeleton-to-video cross-attention for precise control. Together, these components improve visual fidelity, geometric stability, and action alignment in long egocentric rollouts. Moreover, augmenting 400 real trajectories with 400 \method-generated trajectories improves out-of-distribution real-robot success from 77% to 84% on single-arm tasks and from 53% to 70% on dual-arm tasks, demonstrating that the synthesized data substantially improve downstream WAM generalization.
Zexuan Yan, Yuzhou Wu, Yue Ma +9
1Shanghai Jiao Tong University · 2Alibaba Group · 3Tianji KernalMind Co., Ltd. +4
Exo-to-egocentric video generation aims to synthesize what a person sees from their own viewpoint given third-person footage and a target head trajectory. The task requires transferring appearance and semantics across large viewpoint changes while hallucinating content never observed by the exocentric camera. Existing approaches either impose additional input requirements, such as a ground-truth initial egocentric frame or multiple synchronized exocentric views, or remain limited to category-specific settings. EgoX is the first to address cross-activity and in-the-wild generalization, but requires a human-provided caption of the non-existent egocentric view at inference and introduces a computationally expensive geometry-guided attention bias that can propagate reconstruction errors and suppress textual and visual context. We therefore propose \textbf{Ego-Forge}, a caption-free and bias-free framework for exo-to-egocentric generation. It introduces \textit{Dynamic Captioning}, which derives conditioning tokens directly from the model's hidden states and adapts them to the diffusion timestep and network depth, replacing external text conditioning. By scaling training by an order of magnitude and using all available exocentric viewpoints, Ego-Forge learns cross-view correspondence implicitly and eliminates the need for geometry-guided attention, requiring only a lightweight depth prior. Ego-Forge achieves state-of-the-art performance on Ego-Exo4D, runs faster end-to-end, requires no external annotation at inference, and generalizes to in-the-wild scenes, including cases where over-reliance on geometry blocks appearance inference. Our model and source code will be made publicly available.