Generating an egocentric video from a single exocentric recording is a challenging case of novel view synthesis, as the two cameras share little overlap and much of the target view is unobserved. Current state-of-the-art methods reconstruct the scene explicitly by estimating depth, lifting the video into a point cloud, and re-rendering it from the egocentric camera to condition a video diffusion model. This deterministic mapping assigns each pixel to a single reprojected location, which preserves texture but translates depth errors into misplaced content. We ask what a video diffusion model should receive as its condition and propose a lifting-free answer: a learned view synthesizer, an LVSM-style transformer fine-tuned to render the egocentric view directly without depth, point clouds, or reprojection, resolving cross-view correspondence internally. In contrast, its probabilistic mapping averages each region over candidate source locations according to a learned correspondence distribution, preserving structure while fine texture is averaged away. We argue that this trade-off suits a diffusion generator, whose denoising training excels at restoring detail, so an effective condition should prioritize structural alignment over sharpness. This distribution's concentration also yields a per-region confidence, used both to mask low-confidence regions and to guide the generator toward high-confidence areas during early layout-forming denoising steps. Our approach consistently outperforms the state-of-the-art explicit pipeline and generalizes to other datasets without retraining. The synthesizer thus supplies view structure, and the diffusion model its detail.
Figures & tables
Figure 1: Egocentric videos generated by LEGO from real-world exocentric footage, viewed from the perspective of the LEGO minifigure marked by the dotted outline.
Figure 2: Overview of LEGO . The frozen synthesizer F renders the egocentric view once per latent frame and exposes a confidence map. The render, gated to gray where confidence is low, fills the egocentric half of the conditioning latent; the confidence also drives ARC during the early denoising steps. The generator is unchanged.
Figure 3: Condition generation. EgoX estimates depth, lifts the exocentric video into a point cloud, and re-renders it along the egocentric trajectory; LEGO renders the egocentric view directly with a frozen synthesizer.
Image Criteria
Video Criteria
Method
PSNR ↑
SSIM ↑
LPIPS ↓
CLIP-I ↑
FVD ↓
TF ↑
MS ↑
DD ↑
Seen
Exo2Ego-V
14.53
0.384
0.569
0.774
622.47
0.960
0.966
0.985
TrajectoryCrafter
13.05
0.375
0.606
0.780
546.09
0.960
0.980
0.947
Wan Fun Control
12.25
0.463
0.617
0.810
595.07
0.968
0.980
0.901
Wan VACE
12.95
0.413
0.626
0.829
508.69
0.989
0.994
0.673
EgoX
16.05
0.556
0.498
0.896
184.47
0.977
0.990
0.974
Table 1: Comparison on Ego-Exo4D. Unmarked rows are quoted from Kang et al. (2026) ; † released EgoX checkpoint rerun under our scoring stack. TF, MS, and DD are VBench temporal flickering, motion smoothness, and dynamic degree. Best in bold.
Seen
Unseen
Arm
PSNR ↑
LPIPS ↓
CLIP-I ↑
FVD ↓
PSNR ↑
LPIPS ↓
CLIP-I ↑
FVD ↓
Conditioning input: what fills it
No condition
12.30
0.591
0.862
310.00
11.38
0.640
0.838
695.44
Point-cloud render
15.59
0.478
0.895
195.80
13.42
0.595
0.853
549.77
Untuned synthesizer render
15.71
0.503
0.888
212.94
13.56
0.596
0.848
560.51
Synthesizer render
18.24
0.364
0.904
183.40
14.35
0.567
0.860
581.90
Table 2: Ablation under one recipe. Top : what fills the conditioning input (exocentric tokens hidden). Middle : cumulative additions up to the full conditioning. Bottom : how denoising is driven: GGA applied at training and inference, ARC at inference only, and ARC anchored to the point-cloud render with its hole mask as binary confidence. ARC (synthesizer render) is the LEGO row of Table 1 . Best in bold.
Figure 4: Qualitative ablation on one unseen clip, following the rows of Table 2 . The second row shows enlarged crops of the corresponding regions in the first row.
Condition
ZNCC ↑ (patch p×p )
DINO ↑
Sharpness ↑
Time (s) ↓
p=16
32
64
Point-cloud render
0.034
0.047
0.059
0.476
0.97
69.2
Untuned synthesizer render
0.008
0.016
0.030
0.508
0.50
4.7
Synthesizer render
0.169
0.217
0.275
0.586
0.28
4.7
Table 3: Condition quality on the unseen split. ZNCC and DINO co-location cosine measure alignment with the ground truth inside the valid circle; Sharpness is the Laplacian-variance ratio to the ground truth (1 = the recording); Time is seconds per clip on one GPU. Protocol in Appendix B.4 .
Figure 5: Confidence is informative. Windows of the generated video are sorted by confidence; the curve is the error of the most confident fraction relative to all windows, on Ego-Exo4D. Dashed: the average over all windows.
Image Criteria
Video Criteria
Dataset
Method
PSNR ↑
SSIM ↑
LPIPS ↓
CLIP-I ↑
FVD ↓
TF ↑
MS ↑
DD ↑
EgoHumans
EgoX †
13.74
0.460
0.573
0.806
421.83
0.983
0.991
1.000
LEGO
15.52
0.530
0.545
0.819
332.41
0.985
0.992
1.000
Nymeria
EgoX †
11.64
0.451
0.637
0.804
296.98
0.984
0.993
1.000
LEGO
15.23
0.534
0.561
0.821
210.27
0.988
0.994
1.000
Table 4: Evaluation on EgoHumans ( Khirodkar et al., 2023 ) and Nymeria ( Ma et al., 2024 ) with the weights trained on Ego-Exo4D only. Each dataset has 200 test clips. Best in bold.
Figure 6: Qualitative comparison between EgoX and LEGO from the viewpoint indicated by the dotted outline. GT is unavailable (NA) for the in-the-wild clips.
Appendix figures & tables8 assets
Supplementary material from the paper’s appendix.
Appendix
Seen
Unseen
System
LocErr ↓
IoU ↑
Contour ↑
LocErr ↓
IoU ↑
Contour ↑
Fixed protocol (ours; validated on the published unseen row)
EgoX (published)
61.8
0.363
0.546
149.9
0.092
0.481
EgoX (released) †
125.3
0.176
0.622
146.5
0.091
0.572
LEGO
107.7
0.279
0.674
143.3
0.097
0.609
Seen-fitted protocol (closest sweep combination to the published seen row)
Appendix
Table 5: Object-level criteria under two protocols. Top : our fixed protocol, which reproduces the published unseen row from the released checkpoint. Bottom : the sweep combination closest to the published seen row. Published numbers are quoted .
Figure 7: Effect of the gate threshold. The synthesizer render gated at τ=0.1,…,0.5 , shown between the ungated render and the point-cloud render of Kang et al. (2026) .
EgoX †
LEGO
Denoiser bias
GGA
ARC
Denoising loop (s) ↓
565.3
540.3
Generation (s) ↓
597.5
573.2
End-to-end (s) ↓
666.7
577.9
Peak VRAM (GiB) ↓
69.9
68.2
Appendix
Table 6: Cost of generating one clip on one H200 GPU, excluding the one-time model load. Generation assumes that the condition is already built, and end-to-end includes it. Best in bold.
Figure 8: A limitation from F . Where the exocentric frame does not cover the egocentric view, the condition is empty and the generator fills the region from its own prior.
Figure 9: Additional qualitative comparisons on the Ego-Exo4D seen split.
Figure 10: Additional qualitative comparisons on the Ego-Exo4D unseen split.
Figure 11: Additional qualitative comparisons on EgoHumans.
Figure 12: Additional qualitative comparisons on Nymeria.