Generating an egocentric video from a single exocentric recording is a challenging case of novel view synthesis, as the two cameras share little overlap and much of the target view is unobserved. Current state-of-the-art methods reconstruct the scene explicitly by estimating depth, lifting the video into a point cloud, and re-rendering it from the egocentric camera to condition a video diffusion model. This deterministic mapping assigns each pixel to a single reprojected location, which preserves texture but translates depth errors into misplaced content. We ask what a video diffusion model should receive as its condition and propose a lifting-free answer: a learned view synthesizer, an LVSM-style transformer fine-tuned to render the egocentric view directly without depth, point clouds, or reprojection, resolving cross-view correspondence internally. In contrast, its probabilistic mapping averages each region over candidate source locations according to a learned correspondence distribution, preserving structure while fine texture is averaged away. We argue that this trade-off suits a diffusion generator, whose denoising training excels at restoring detail, so an effective condition should prioritize structural alignment over sharpness. This distribution's concentration also yields a per-region confidence, used both to mask low-confidence regions and to guide the generator toward high-confidence areas during early layout-forming denoising steps. Our approach consistently outperforms the state-of-the-art explicit pipeline and generalizes to other datasets without retraining. The synthesizer thus supplies view structure, and the diffusion model its detail.
Figures & tables
Figure 1: Egocentric videos generated by LEGO from real-world exocentric footage, viewed from the perspective of the LEGO minifigure marked by the dotted outline.
Figure 2: Overview of LEGO . The frozen synthesizer F renders the egocentric view once per latent frame and exposes a confidence map. The render, gated to gray where confidence is low, fills the egocentric half of the conditioning latent; the confidence also drives ARC during the early denoising steps. The generator is unchanged.
Figure 3: Condition generation. EgoX estimates depth, lifts the exocentric video into a point cloud, and re-renders it along the egocentric trajectory; LEGO renders the egocentric view directly with a frozen synthesizer.
Image Criteria
Video Criteria
Method
PSNR ↑
SSIM ↑
LPIPS ↓
CLIP-I ↑
FVD ↓
TF ↑
MS ↑
DD ↑
Seen
Exo2Ego-V
14.53
0.384
0.569
0.774
622.47
0.960
0.966
0.985
TrajectoryCrafter
13.05
0.375
0.606
0.780
546.09
0.960
0.980
0.947
Wan Fun Control
12.25
0.463
0.617
0.810
595.07
0.968
0.980
0.901
Wan VACE
12.95
0.413
0.626
0.829
508.69
0.989
0.994
0.673
EgoX
16.05
0.556
0.498
0.896
184.47
0.977
0.990
0.974
Table 1: Comparison on Ego-Exo4D. Unmarked rows are quoted from Kang et al. (2026) ; † released EgoX checkpoint rerun under our scoring stack. TF, MS, and DD are VBench temporal flickering, motion smoothness, and dynamic degree. Best in bold.
Seen
Unseen
Arm
PSNR ↑
LPIPS ↓
CLIP-I ↑
FVD ↓
PSNR ↑
LPIPS ↓
CLIP-I ↑
FVD ↓
Conditioning input: what fills it
No condition
12.30
0.591
0.862
310.00
11.38
0.640
0.838
695.44
Point-cloud render
15.59
0.478
0.895
195.80
13.42
0.595
0.853
549.77
Untuned synthesizer render
15.71
0.503
0.888
212.94
13.56
0.596
0.848
560.51
Synthesizer render
18.24
0.364
0.904
183.40
14.35
0.567
0.860
581.90
Table 2: Ablation under one recipe. Top : what fills the conditioning input (exocentric tokens hidden). Middle : cumulative additions up to the full conditioning. Bottom : how denoising is driven: GGA applied at training and inference, ARC at inference only, and ARC anchored to the point-cloud render with its hole mask as binary confidence. ARC (synthesizer render) is the LEGO row of Table 1 . Best in bold.
Figure 4: Qualitative ablation on one unseen clip, following the rows of Table 2 . The second row shows enlarged crops of the corresponding regions in the first row.
Condition
ZNCC ↑ (patch p×p )
DINO ↑
Sharpness ↑
Time (s) ↓
p=16
32
64
Point-cloud render
0.034
0.047
0.059
0.476
0.97
69.2
Untuned synthesizer render
0.008
0.016
0.030
0.508
0.50
4.7
Synthesizer render
0.169
0.217
0.275
0.586
0.28
4.7
Table 3: Condition quality on the unseen split. ZNCC and DINO co-location cosine measure alignment with the ground truth inside the valid circle; Sharpness is the Laplacian-variance ratio to the ground truth (1 = the recording); Time is seconds per clip on one GPU. Protocol in Appendix B.4 .
Figure 5: Confidence is informative. Windows of the generated video are sorted by confidence; the curve is the error of the most confident fraction relative to all windows, on Ego-Exo4D. Dashed: the average over all windows.
Image Criteria
Video Criteria
Dataset
Method
PSNR ↑
SSIM ↑
LPIPS ↓
CLIP-I ↑
FVD ↓
TF ↑
MS ↑
DD ↑
EgoHumans
EgoX †
13.74
0.460
0.573
0.806
421.83
0.983
0.991
1.000
LEGO
15.52
0.530
0.545
0.819
332.41
0.985
0.992
1.000
Nymeria
EgoX †
11.64
0.451
0.637
0.804
296.98
0.984
0.993
1.000
LEGO
15.23
0.534
0.561
0.821
210.27
0.988
0.994
1.000
Table 4: Evaluation on EgoHumans ( Khirodkar et al., 2023 ) and Nymeria ( Ma et al., 2024 ) with the weights trained on Ego-Exo4D only. Each dataset has 200 test clips. Best in bold.
Figure 6: Qualitative comparison between EgoX and LEGO from the viewpoint indicated by the dotted outline. GT is unavailable (NA) for the in-the-wild clips.
Appendix figures & tables8 assets
Supplementary material from the paper’s appendix.
Appendix
Seen
Unseen
System
LocErr ↓
IoU ↑
Contour ↑
LocErr ↓
IoU ↑
Contour ↑
Fixed protocol (ours; validated on the published unseen row)
EgoX (published)
61.8
0.363
0.546
149.9
0.092
0.481
EgoX (released) †
125.3
0.176
0.622
146.5
0.091
0.572
LEGO
107.7
0.279
0.674
143.3
0.097
0.609
Seen-fitted protocol (closest sweep combination to the published seen row)
Appendix
Table 5: Object-level criteria under two protocols. Top : our fixed protocol, which reproduces the published unseen row from the released checkpoint. Bottom : the sweep combination closest to the published seen row. Published numbers are quoted .
Figure 7: Effect of the gate threshold. The synthesizer render gated at τ=0.1,…,0.5 , shown between the ungated render and the point-cloud render of Kang et al. (2026) .
EgoX †
LEGO
Denoiser bias
GGA
ARC
Denoising loop (s) ↓
565.3
540.3
Generation (s) ↓
597.5
573.2
End-to-end (s) ↓
666.7
577.9
Peak VRAM (GiB) ↓
69.9
68.2
Appendix
Table 6: Cost of generating one clip on one H200 GPU, excluding the one-time model load. Generation assumes that the condition is already built, and end-to-end includes it. Best in bold.
Figure 8: A limitation from F . Where the exocentric frame does not cover the egocentric view, the condition is empty and the generator fills the region from its own prior.
Figure 9: Additional qualitative comparisons on the Ego-Exo4D seen split.
Figure 10: Additional qualitative comparisons on the Ego-Exo4D unseen split.
Figure 11: Additional qualitative comparisons on EgoHumans.
Figure 12: Additional qualitative comparisons on Nymeria.
Exo-to-egocentric video generation aims to synthesize what a person sees from their own viewpoint given third-person footage and a target head trajectory. The task requires transferring appearance and semantics across large viewpoint changes while hallucinating content never observed by the exocentric camera. Existing approaches either impose additional input requirements, such as a ground-truth initial egocentric frame or multiple synchronized exocentric views, or remain limited to category-specific settings. EgoX is the first to address cross-activity and in-the-wild generalization, but requires a human-provided caption of the non-existent egocentric view at inference and introduces a computationally expensive geometry-guided attention bias that can propagate reconstruction errors and suppress textual and visual context. We therefore propose \textbf{Ego-Forge}, a caption-free and bias-free framework for exo-to-egocentric generation. It introduces \textit{Dynamic Captioning}, which derives conditioning tokens directly from the model's hidden states and adapts them to the diffusion timestep and network depth, replacing external text conditioning. By scaling training by an order of magnitude and using all available exocentric viewpoints, Ego-Forge learns cross-view correspondence implicitly and eliminates the need for geometry-guided attention, requiring only a lightweight depth prior. Ego-Forge achieves state-of-the-art performance on Ego-Exo4D, runs faster end-to-end, requires no external annotation at inference, and generalizes to in-the-wild scenes, including cases where over-reliance on geometry blocks appearance inference. Our model and source code will be made publicly available.
Foundation video generation models such as WAN 2.2 exhibit strong text- and image-conditioned synthesis abilities but remain constrained to the same-view generation setting. In this work, we introduce Exo2EgoSyn, an adaptation of WAN 2.2 that unlocks Exocentric-to-Egocentric(Exo2Ego) cross-view video synthesis. Our framework consists of three key modules. Ego-Exo View Alignment(EgoExo-Align) enforces latent-space alignment between exocentric and egocentric first-frame representations, reorienting the generative space from the given exo view toward the ego view. Multi-view Exocentric Video Conditioning (MultiExoCon) aggregates multi-view exocentric videos into a unified conditioning signal, extending WAN2.2 beyond its vanilla single-image or text conditioning. Furthermore, Pose-Aware Latent Injection (PoseInj) injects relative exo-to-ego camera pose information into the latent state, guiding geometry-aware synthesis across viewpoints. Together, these modules enable high-fidelity ego view video generation from third-person observations without retraining from scratch. Experiments on ExoEgo4D validate that Exo2EgoSyn significantly improves Ego2Exo synthesis, paving the way for scalable cross-view video generation with foundation models. Source code and models will be released publicly.
Recent advances in video world models enable interactive environments with free navigation, making translation between first-person (egocentric) and third-person (exocentric) perspectives increasingly important. However, existing studies focus on unidirectional exocentric-to-egocentric translation, overlooking reference-guided exocentric perspective synthesis. This capability is crucial for gaming and embodied AI applications. Motivated by this, we present WorldWander, an in-context learning framework tailored for translating between egocentric and exocentric worlds in video generation. Building upon advanced video diffusion transformers, WorldWander integrates (i) In-Context Perspective Alignment and (ii) Collaborative Position Encoding to model cross-view synchronization and character consistency. To support our task, we curate EgoExo-8K, a dynamic and scene-rich dataset containing synchronized egocentric-exocentric triplets from both synthetic and real-world scenarios. Experiments demonstrate that WorldWander achieves superior perspective synchronization, character consistency, and generalization, setting a new benchmark for egocentric-exocentric video translation.
Quanjian Song, Yiren Song, Kelly Peng +2
First Intelligence · Show Lab, National University of Singapore, Singapore