cs.CVOct 8, 2026

LEGO: A Lifting-Free Approach for Exocentric-to-Egocentric Video Generation

Authors: Suhwan Cho, Yonwoo Choi, Soongjin Kim, Jicheol Park, Taegyu Lim

Organizations: GenGenAI

Abstract

Generating an egocentric video from a single exocentric recording is a challenging case of novel view synthesis, as the two cameras share little overlap and much of the target view is unobserved. Current state-of-the-art methods reconstruct the scene explicitly by estimating depth, lifting the video into a point cloud, and re-rendering it from the egocentric camera to condition a video diffusion model. This deterministic mapping assigns each pixel to a single reprojected location, which preserves texture but translates depth errors into misplaced content. We ask what a video diffusion model should receive as its condition and propose a lifting-free answer: a learned view synthesizer, an LVSM-style transformer fine-tuned to render the egocentric view directly without depth, point clouds, or reprojection, resolving cross-view correspondence internally. In contrast, its probabilistic mapping averages each region over candidate source locations according to a learned correspondence distribution, preserving structure while fine texture is averaged away. We argue that this trade-off suits a diffusion generator, whose denoising training excels at restoring detail, so an effective condition should prioritize structural alignment over sharpness. This distribution's concentration also yields a per-region confidence, used both to mask low-confidence regions and to guide the generator toward high-confidence areas during early layout-forming denoising steps. Our approach consistently outperforms the state-of-the-art explicit pipeline and generalizes to other datasets without retraining. The synthesizer thus supplies view structure, and the diffusion model its detail.

Figures & tables

Appendix figures & tables8 assets

Supplementary material from the paper’s appendix.

Appendix

Explore similar work

CardsList
  1. Ego-Forge: Text and Geometric-Attention Free Exo-to-Egocentric Video Generation

    Sep 28, 2026Mohammad Mahdi, Luc Van Gool, Danda Pani PaudelCamera-Conditioned Video GenerationVideo Diffusion Models

  2. Exo2EgoSyn: Unlocking Foundation Video Generation Models for Exocentric-to-Egocentric Video Synthesis

    Nov 25, 2025Mohammad Mahdi, Yuqian Fu, Nedko Savov +3Camera-Conditioned Video GenerationVideo Diffusion Models

  3. WorldWander: Bridging Egocentric and Exocentric Worlds in Video Generation

    Nov 27, 2025Quanjian Song, Yiren Song, Kelly Peng +2Video Diffusion ModelsVideo World Models