Exo-to-egocentric video generation aims to synthesize what a person sees from their own viewpoint given third-person footage and a target head trajectory. The task requires transferring appearance and semantics across large viewpoint changes while hallucinating content never observed by the exocentric camera. Existing approaches either impose additional input requirements, such as a ground-truth initial egocentric frame or multiple synchronized exocentric views, or remain limited to category-specific settings. EgoX is the first to address cross-activity and in-the-wild generalization, but requires a human-provided caption of the non-existent egocentric view at inference and introduces a computationally expensive geometry-guided attention bias that can propagate reconstruction errors and suppress textual and visual context. We therefore propose \textbf{Ego-Forge}, a caption-free and bias-free framework for exo-to-egocentric generation. It introduces \textit{Dynamic Captioning}, which derives conditioning tokens directly from the model's hidden states and adapts them to the diffusion timestep and network depth, replacing external text conditioning. By scaling training by an order of magnitude and using all available exocentric viewpoints, Ego-Forge learns cross-view correspondence implicitly and eliminates the need for geometry-guided attention, requiring only a lightweight depth prior. Ego-Forge achieves state-of-the-art performance on Ego-Exo4D, runs faster end-to-end, requires no external annotation at inference, and generalizes to in-the-wild scenes, including cases where over-reliance on geometry blocks appearance inference. Our model and source code will be made publicly available.
Figures & tables
Figure 1: Ego-Forge generates egocentric video without text conditioning or geometry-guided attention. Top: An in-the-wild exocentric clip in which the subject turns away from a mirror and then back. Middle: Our method follows the head motion and renders the mirror and its reflection, although the exocentric camera never observes the reflected content and no caption is provided at inference. Bottom: EgoX, conditioned on captions of both views, including an explicit description of the reflection, and Geometry-Guided Self-Attention (GGA), fails to recover the reflection.
Figure 2: Overview of Ego-Forge. Throughout, S denotes self-attention and X cross-attention. (1) Text-conditioned adaptation. We adapt the pretrained video diffusion backbone using all available exocentric viewpoints, conditioned on the exocentric video, reprojected egocentric prior, and a textual caption, without geometric attention bias. (2) Learning Dynamic Captioning. A single resampler R , shared across all N blocks, learns to replace the textual conditioning using the hidden state hi , block index i , and diffusion timestep t . Here, the superscripts t and s denote the teacher and student, respectively. (3) Joint finetuning. The caption is removed and the model is further finetuned with Dynamic Captioning and self-attention under the diffusion objective. (4) Inference. Only the exocentric video and reprojected prior are required; R provides the conditioning throughout the network without a caption.
Figure 3: R’s attention structure. Mx and Mg apply Eq 4 . The black block masks exo queries from ego keys.
Figure 4: Qualitative results. Left: a) Ours better positions the arm on the man’s chest. b) Ours handles small exo-camera motion. c) 1 : Ours uses visual cues, e.g., the exo-view shadow. 2 : Ours better synthesizes unseen content, using the car observed in the exo view. 3 : The baseline misinterprets the “EMERGENCY” sign (mentioned in caption) and generates it at the wrong time.
Image Criteria
Video Criteria
Scenes
Method
PSNR ↑
SSIM ↑
LPIPS ↓
CLIP-I ↑
FVD ↓
Temporal Flickering ↑
Motion Smoothness ↑
Dynamic Degree ↑
Seen
Exo2Ego-V
14.41
0.382
0.564
0.794
642.09
0.953
0.944
0.986
Trj-Crafter
13.59
0.397
0.591
0.788
741.72
0.960
0.980
0.950
Wan-FCtrl
13.11
0.433
0.622
0.789
610.13
0.966
0.980
0.906
Wan VACE
13.48
0.453
0.611
0.771
583.29
0.989
0.994
0.691
Syn2Seq-F
15.01
0.461
0.549
0.793
513.73
0.965
0.970
0.822
Table 1: Quantitative comparison on Ego-Exo4D. Bold and underlined denote the best and second-best results, respectively. Trj-Crafter ( Yu et al., 2025 ) , Wan-FCtrl ( AIGC-Apps, 2024 ) , Wan VACE ( Jiang et al., 2025 ) , and Syn2Seq-F ( Mahdi et al., 2026 ) .
Figure 6Figure 7Figure 8Figure 9
Appendix figures & tables2 assets
Supplementary material from the paper’s appendix.
Appendix
Figure A: EgoX ( Kang et al., 2026 ) performance depending on GGA presence.
Stage
GPUs
Days
Trainable
1: Text-conditioned adaptation
2
5
LoRA on S , X ; patch embedding
2: Learning the Dynamic Captioner
1
3
R only
3: Caption-free finetuning
4
2
R and LoRA on S
Appendix
Table A: Training budget. All stages run at batch size 1 per GPU with gradient checkpointing; the 14 B backbone and the stitched canvas leave little memory for a larger batch.
Foundation video generation models such as WAN 2.2 exhibit strong text- and image-conditioned synthesis abilities but remain constrained to the same-view generation setting. In this work, we introduce Exo2EgoSyn, an adaptation of WAN 2.2 that unlocks Exocentric-to-Egocentric(Exo2Ego) cross-view video synthesis. Our framework consists of three key modules. Ego-Exo View Alignment(EgoExo-Align) enforces latent-space alignment between exocentric and egocentric first-frame representations, reorienting the generative space from the given exo view toward the ego view. Multi-view Exocentric Video Conditioning (MultiExoCon) aggregates multi-view exocentric videos into a unified conditioning signal, extending WAN2.2 beyond its vanilla single-image or text conditioning. Furthermore, Pose-Aware Latent Injection (PoseInj) injects relative exo-to-ego camera pose information into the latent state, guiding geometry-aware synthesis across viewpoints. Together, these modules enable high-fidelity ego view video generation from third-person observations without retraining from scratch. Experiments on ExoEgo4D validate that Exo2EgoSyn significantly improves Ego2Exo synthesis, paving the way for scalable cross-view video generation with foundation models. Source code and models will be released publicly.
Exo-to-Ego video generation aims to synthesize a first-person video from a synchronized third-person view and corresponding camera poses. While paired supervision is available, synchronized exo-ego data inherently introduces substantial spatio-temporal and geometric discontinuities, violating the smooth-motion assumptions of standard video generation benchmarks. We identify this synchronization-induced jump as the central challenge and propose Syn2Seq-Forcing, a sequential formulation that interpolates between the source and target videos to form a single continuous signal. By reframing Exo2Ego as sequential signal modeling rather than a conventional condition-output task, our approach enables diffusion-based sequence models, e.g. Diffusion Forcing Transformers (DFoT), to capture coherent transitions across frames more effectively. Empirically, we show that interpolating only the videos, without performing pose interpolation already produces significant improvements, emphasizing that the dominant difficulty arises from spatio-temporal discontinuities. Beyond immediate performance gains, this formulation establishes a general and flexible framework capable of unifying both Exo2Ego and Ego2Exo generation within a single continuous sequence model, providing a principled foundation for future research in cross-view video synthesis.
Egocentric videos of human manipulation provide valuable visual experience for embodied intelligence, yet collecting such data at scale is costly. Exocentric-to-egocentric video generation offers a scalable alternative by transforming abundant third-person manipulation videos into first-person observations. However, existing methods often struggle to faithfully preserve demonstrated hand-object interactions (HOI) across large viewpoint changes due to insufficient fine-grained interaction guidance and weak object-centric anchoring. We present Exo2EgoHOI, an HOI-aware video generative framework for interaction-preserving exocentric-to-egocentric translation. To preserve fine-grained HOI, we introduce a unified 4D HOI prior that combines scene geometry, articulated hand renderings, and dense hand-object relation fields, together with a dual-branch residual adapter for injecting structural and relational cues into the video generation backbone. To preserve object consistency, we introduce Decomposed Gated Cross-Attention, which separately encodes object and background references and adaptively integrates global semantic and local appearance features as object-centric anchors. Experiments on ARCTIC-HOI and Ego-Exo4D demonstrate substantial improvements in object consistency and HOI preservation while maintaining competitive visual fidelity. In particular, on ARCTIC-HOI, Exo2EgoHOI improves object mIoU by 32.3% and reduces MPJPE and PA-MPJPE by 34.7% and 50.0%, respectively, relative to the respective best baseline results. Project page: https://rcl-robotics.github.io/Exo2EgoHOI/.