Generative world models have shown promise for simulating dynamic environments, yet egocentric video remains challenging due to rapid viewpoint changes, frequent hand-object interactions, and goal-directed procedures whose evolution depends on latent human intent. Existing approaches either focus on hand-centric instructional synthesis with limited scene evolution, perform static view translation without modeling action dynamics, or rely on dense supervision, such as camera trajectories, long video prefixes, and synchronized multi-camera capture. In this work, we introduce EgoForge, an egocentric goal-directed world simulator that generates coherent, first-person video rollouts from minimal static inputs: a single egocentric image, a high-level instruction, and an optional auxiliary exocentric view. To improve intent alignment and temporal coherence, we introduce GRAFT, a trajectory-level diffusion refinement method that uses positive and negative rollout distributions, derived from goal, temporal, scene-consistency, and perceptual rewards, to steer the diffusion velocity field toward coherent, goal-complete egocentric simulations. Extensive experiments show EgoForge achieves consistent gains in semantic alignment, geometric stability, and motion fidelity over strong baselines, and performs robustly in real-world smart-glasses experiments.
Figures & tables
Figure 1 : EgoForge Overview : Given a single egocentric observation, a high-level instruction expressing user intent, and an auxiliary exo-view reference, EgoForge fuses encoded visual features with noisy video latents at each DiT block to guide generation. Geometry alignment weakly supervises intermediate features using angular and scale consistency to encourage spatially stable rollouts. The resulting rollout videos are further refined via the proposed GRAFT alignment that optimizes goal completion, scene consistency, temporal causality, and perceptual fidelity.
Figure 2 : GRAFT refinement. Rollout-level rewards yield normalized weights r and 1−r for the implicit positive and negative branch losses. Only the trainable model receives gradients, while the old model is refreshed across iterations.
Model
DINO-Score ↑
CLIP-Score ↑
SSIM ↑
LPIPS ↓
FVD ↓
flow MSE ↓
PSNR ↑
EgoDreamer [ 61 ]
42.35
25.40
0.58
0.35
580.45
8.15
15.20
Handi [ 31 ]
31.12
18.25
0.42
0.52
912.30
14.50
12.85
Cosmos [ 42 ]
49.42
29.77
0.70
0.26
448.12
6.40
18.73
HunyuanVideo [ 28 ]
53.54
29.43
0.71
0.26
384.31
6.10
18.88
WAN2.2 [ 58 ]
53.99
35.69
0.72
0.23
322.17
5.78
20.44
EgoForge (Ours)
61.25
39.30
0.79
0.15
182.25
2.83
24.08
Table 2 : Quantitative comparisons on the X-Ego benchmark. EgoForge outperforms all baselines across semantic, perceptual, and temporal metrics.
Model
DINO-Score ↑
CLIP-Score ↑
SSIM ↑
LPIPS ↓
FVD ↓
flow MSE ↓
PSNR ↑
Cosmos+EV
48.60
29.60
0.67
0.28
485.75
6.82
18.30
Cosmos+TT
50.80
30.40
0.71
0.25
433.90
6.31
18.88
HunyuanVideo+EV
52.80
29.20
0.70
0.27
405.87
6.30
18.61
HunyuanVideo+TT
54.10
29.86
0.72
0.24
365.80
5.95
19.10
WAN2.2+EV
52.91
35.11
0.71
0.27
352.41
6.25
20.05
WAN2.2+TT
54.80
36.20
0.73
0.25
310.57
5.60
20.64
Table 3 : Quantitative comparisons on X-Ego between EgoForge and other finetuned baseline variants. +EV: with exo-view img, +TT: text-only domain adaptation; +CI: our conditioning inputs (exo-view and goal instructions) with our structured injection with Geometry Weak Supervision.
FT
GWS
GRAFT
DINO-Score ↑
CLIP-Score ↑
SSIM ↑
LPIPS ↓
FVD ↓
flow MSE ↓
PSNR ↑
✓
✗
✗
56.81
37.10
0.74
0.21
260.89
4.82
21.92
✓
✓
✗
58.92
38.05
0.76
0.18
218.72
3.92
22.87
✓
✓
✓
61.25
39.30
0.79
0.15
182.25
2.83
24.08
Table 4 : Ablation on EgoForge modules. We evaluate the impact of diffusion finetuning (FT), geometry weak supervision (GWS), and the proposed GRAFT refinement. Each component consistently improves performance across semantic, perceptual, and temporal metrics, with the full EgoForge model achieving the best overall results.
Conditioning
Inputs
Generation Quality
Ego
Text
Exo
DINO ↑
CLIP ↑
SSIM ↑
LPIPS ↓
FVD ↓
Flow MSE ↓
PSNR ↑
Single-modality conditioning
Text only
✗
✓
✗
49.84
36.75
0.62
0.33
398.52
7.01
17.18
Ego only
✓
✗
✗
54.26
28.41
0.70
0.25
310.44
5.73
20.14
Exo only
✗
✗
✓
50.91
27.86
0.65
0.30
365.73
6.38
18.56
Two-modality conditioning
Table 5 : Ablation of conditioning inputs. Missing modalities are replaced with learned null embeddings. For counterfactual conditions, random Exo uses an exocentric view sampled from another test example, and wrong Text uses an action-incompatible instruction. Metrics computed against original held-out video and instruction.
Figure 3 : Qualitative Comparison between EgoForge and baselines . EgoForge accurately reconstructs multi-step, causally ordered actions, preserving hand–object geometry, temporal consistency, and goal alignment. For instance, in the first example, Cosmos erroneously generates a third hand, Hunyuan depicts a disconnected arm, and Wan2.2 fails to complete the coffee-pouring task. In the second example, Cosmos generates multiple balls, while both Hunyuan and Wan2.2 generate an incorrect person to perform the action. In contrast, EgoForge accurately completes both tasks.
Figure 4 : Qualitative comparison with and without exocentric input. EgoForge can be steered using an auxiliary exo-view to improve spatial grounding by anchoring the simulation to the reference environment.
Appendix figures & tables5 assets
Supplementary material from the paper’s appendix.
Appendix
Rewards
DINO-Score ↑
CLIP-Score ↑
SSIM ↑
LPIPS ↓
FVD ↓
flow MSE ↓
PSNR ↑
✗ \penalty\penaltyRgoal
59.62
38.49
0.78
0.16
205.96
3.48
23.48
✗ \penalty\penaltyRenv
60.67
39.05
0.78
0.16
200.49
3.43
23.60
✗ \penalty\penaltyRtemp
60.78
39.11
0.78
0.16
213.25
3.70
23.72
✗ \penalty\penaltyRper
60.32
38.80
0.77
0.18
204.13
3.48
23.17
EgoForge (Ours)
61.25
39.30
0.79
0.15
182.25
2.83
24.08
Appendix
Table 16 : Effect of reward components in GRAFT . Each reward term contributes to improved semantic alignment and temporal coherence, while the full reward composition yields the strongest performance across all metrics.
Figure 5 : Qualitative egocentric video rollouts. EgoForge generates temporally coherent first-person video trajectories that follow the intended activity while preserving scene structure and realistic hand–object interactions across diverse environments.
Figure 6 : Qualitative Comparison between EgoForge and baselines . Sample frames from generated videos illustrating two scenarios. Top row: In the hand-washing task, baselines struggle with object consistency ( e.g. , Cosmos hallucinates the soap source) or ignore scene context ( e.g. , Wan2.2 bbypasses the soap on the table), while EgoForge successfully executes the action. using existing objects. Bottom row: In the soccer task, baselines exhibit severe artifacts like ghosting (Cosmos) or fail to follow precise instructions regarding motion and goals (Hunyuan, Wan2.2). EgoForge accurately executes the complex command: trapping with the left leg and shooting with the right.
Figure 7 : Dense visualization of long-duration sequences. For each EgoForge-generated video, we show 26 uniformly sampled frames to illustrate how actions unfold over time. Across diverse first-person tasks, EgoForge preserves scene structure, maintains a stable egocentric viewpoint, and produces smooth, goal-consistent hand–object interactions.