Egocentric videos of human manipulation provide valuable visual experience for embodied intelligence, yet collecting such data at scale is costly. Exocentric-to-egocentric video generation offers a scalable alternative by transforming abundant third-person manipulation videos into first-person observations. However, existing methods often struggle to faithfully preserve demonstrated hand-object interactions (HOI) across large viewpoint changes due to insufficient fine-grained interaction guidance and weak object-centric anchoring. We present Exo2EgoHOI, an HOI-aware video generative framework for interaction-preserving exocentric-to-egocentric translation. To preserve fine-grained HOI, we introduce a unified 4D HOI prior that combines scene geometry, articulated hand renderings, and dense hand-object relation fields, together with a dual-branch residual adapter for injecting structural and relational cues into the video generation backbone. To preserve object consistency, we introduce Decomposed Gated Cross-Attention, which separately encodes object and background references and adaptively integrates global semantic and local appearance features as object-centric anchors. Experiments on ARCTIC-HOI and Ego-Exo4D demonstrate substantial improvements in object consistency and HOI preservation while maintaining competitive visual fidelity. In particular, on ARCTIC-HOI, Exo2EgoHOI improves object mIoU by 32.3% and reduces MPJPE and PA-MPJPE by 34.7% and 50.0%, respectively, relative to the respective best baseline results. Project page: https://rcl-robotics.github.io/Exo2EgoHOI/.
Figures & tables
Figure 1: From exocentric human demonstrations to egocentric observations and robot execution. Exo2EgoHOI preserves hand-object interactions during exocentric-to-egocentric video generation, enabling downstream robot execution via retargeting.
Figure 2: Overview of Exo2EgoHOI. Given an exocentric video, we reconstruct world-space 4D HOI and transform it into a unified egocentric prior encoding scene geometry, hand motion, and interaction cues. HOI adapters inject this prior into a video Diffusion Transformer (DiT), while Decomposed Gated Cross-Attention anchors object-specific semantic and appearance cues from the exocentric reference to preserve object consistency throughout egocentric video generation.
Figure 3: Architecture of DiT block. DGCA adaptively fuses global and local contexts through learnable block-specific gates.
Visual Fidelity
Object Consistency
Hand Consistency
Method
PSNR ↑
SSIM ↑
LPIPS ↓
CLIP-I ↑
mIoU ↑
Center-Err. ↓
CLIP-O ↑
MPJPE ↓
WA-MPJPE ↓
PA-MPJPE ↓
ARCTIC-HOI
WAN VACE ( Jiang et al., 2025 )
12.78
0.6163
0.5115
0.8485
0.2336
0.0893
0.8080
2870.22
371.18
20.36
EgoWorld ( Park et al., 2026 )
11.90
0.4652
0.6203
0.7995
0.2241
0.1131
0.7869
2282.03
440.60
57.92
Vista4D ( Lin et al., 2026b )
13.58
0.6614
0.4324
0.8737
0.2768
0.0929
0.8176
2123.78
387.73
28.91
EgoX ( Kang et al., 2026 )
11.83
0.5941
0.4743
0.8910
0.2933
0.0804
0.8345
2019.81
325.56
22.12
Table 1: Quantitative comparison on ARCTIC-HOI and Ego-Exo4D datasets. We evaluate visual fidelity, object consistency, and hand-pose consistency. Best results are shown in bold and second-best results are underlined .
Methods
HOI Priors
DGCA
Evaluation Metrics
Scene
Hand
Inter.
Global
Local
Gate
LPIPS ↓
CLIP-I ↑
mIoU ↑
CLIP-O ↑
MPJPE ↓
WA-MPJPE ↓
PA-MPJPE ↓
#0 Baseline
✗
✗
✗
✗
✗
✗
0.5577
0.8253
0.1148
0.8056
2631.66
371.25
41.42
#1 w/o All Priors
✗
✗
✗
✓
✓
✓
0.5351
0.8449
0.1378
0.8255
2254.26
346.98
21.11
#2 w/ Scene
✓
✗
✗
✓
✓
✓
0.4435
0.9008
0.3104
0.8391
1978.57
327.60
14.30
#3 w/ Scene & Hand
✓
✓
✗
✓
✓
✓
0.4063
0.9108
0.3323
0.8421
1570.51
315.28
13.33
#4 w/o Decomp. CA
✓
✓
✓
✗
✗
✗
0.3953
0.8853
0.3023
0.8198
2327.34
398.94
16.02
Table 2: Ablation studies on the ARCTIC-HOI dataset. “#8 Vanilla CA” uses CLIP to encode the entire anchor frame for cross-attention without decomposition. Best results are shown in bold and second-best results are underlined .
Figure 4: Qualitative comparison on ARCTIC-HOI dataset. Compared with existing methods, Exo2EgoHOI better preserves the target-view object appearance and hand-object configuration under large viewpoint changes. Insets visualize hand meshes recovered from the generated egocentric frames for comparison with the reference.
Figure 5: Ablation studies on the ARCTIC-HOI dataset. Left: qualitative comparison of different configurations. Right: visual fidelity (PSNR and LPIPS; top) and hand/object consistency (PA-MPJPE and mIoU; bottom) with different numbers of exocentric training views.
Appendix figures & tables5 assets
Supplementary material from the paper’s appendix.
Appendix
Operation
Channels
Kernel
Stride
Output (T,H,W)
Branch Conv 1–3
C→16→16→16
(3,3,3)
(1,1,1)
(49,448,448)
Branch Conv 4
16→16
(3,3,3)
(1,2,2)
(49,224,224)
Branch Conv 5
16→16
(3,3,3)
(2,2,2)
(25,112,112)
Branch Conv 6
16→16
(3,3,3)
(2,2,2)
(13,56,56)
Concatenate + fuse
16+16→16
(1,1,1)
(1,1,1)
(13,56,56)
Patch projection
16→5120
(1,2,2)
(1,2,2)
(13,28,28)
Appendix
Table 3: HOI adapter layers; C=3 for hand mesh prior and C=5 for interaction prior. Branch convolutions use padding one and SiLU, fusion and projection use zero padding.
Figure 6: Left: reconstructed 3D scene and camera configuration, with the corresponding egocentric observation shown below. Colored camera frustums denote different exocentric viewpoints. Right: corresponding images of the same scene observed from these exocentric cameras.
Figure 7: Additional qualitative comparison on ARCTIC-HOI. We compare egocentric videos generated by different methods and overlay their reconstructed 3D hand meshes. Hand-mesh vertices are color-coded by errors relative to the egocentric ground truth, ranging from blue (low error) to red (high error); N/A indicates failed hand reconstruction. The visualization highlights differences in target-view hand configuration and hand-object alignment across methods.
Figure 8: Additional qualitative comparison on Ego-Exo4D. We compare egocentric videos generated from exocentric inputs across diverse real-world scenes and manipulation activities. The results illustrate differences among methods in cross-view scene reconstruction, manipulated-object consistency, and hand-object configurations under substantial viewpoint changes.
Figure 9: Representative failure cases of Exo2EgoHOI. Severe occlusion can lead to incomplete hand-object reconstruction, while complex articulated or thin-structured objects remain difficult to model accurately. Insets highlight discrepancies in object geometry, articulation state, and hand-object configuration compared with the egocentric ground truth.
Exo-to-egocentric video generation aims to synthesize what a person sees from their own viewpoint given third-person footage and a target head trajectory. The task requires transferring appearance and semantics across large viewpoint changes while hallucinating content never observed by the exocentric camera. Existing approaches either impose additional input requirements, such as a ground-truth initial egocentric frame or multiple synchronized exocentric views, or remain limited to category-specific settings. EgoX is the first to address cross-activity and in-the-wild generalization, but requires a human-provided caption of the non-existent egocentric view at inference and introduces a computationally expensive geometry-guided attention bias that can propagate reconstruction errors and suppress textual and visual context. We therefore propose \textbf{Ego-Forge}, a caption-free and bias-free framework for exo-to-egocentric generation. It introduces \textit{Dynamic Captioning}, which derives conditioning tokens directly from the model's hidden states and adapts them to the diffusion timestep and network depth, replacing external text conditioning. By scaling training by an order of magnitude and using all available exocentric viewpoints, Ego-Forge learns cross-view correspondence implicitly and eliminates the need for geometry-guided attention, requiring only a lightweight depth prior. Ego-Forge achieves state-of-the-art performance on Ego-Exo4D, runs faster end-to-end, requires no external annotation at inference, and generalizes to in-the-wild scenes, including cases where over-reliance on geometry blocks appearance inference. Our model and source code will be made publicly available.
Controllable video generation for complex hand-object interactions is a critical step toward building visual world models. However, existing methods often struggle to achieve fine-grained, 3D-consistent hand articulation in generated videos. By relying on dense 2D trajectories or implicit pose representations, they collapse crucial geometric structures into spatially ambiguous signals, leading to severe motion inconsistencies and hallucinated artifacts under egocentric occlusions. To address this, we propose leveraging sparse 3D hand joints as explicit control signals with three key advantages: explicit geometry to resolve occlusions, an intuitive interface for interactive editing, and cross-embodiment generalization to robotic hands. Built upon this, our efficient control module extracts occlusion-aware features from the source reference frame by penalizing unreliable visual features from hidden joints, and employs a 3D-based weighting mechanism to handle dynamically occluded target joints during motion propagation. Meanwhile, it directly injects 3D geometric embeddings into the latent space to enforce structural consistency. To facilitate robust training and evaluation, we develop an automated annotation pipeline, yielding 1M high-quality egocentric video clips paired with precise hand trajectories. Experiments demonstrate that our approach outperforms state-of-the-art baselines, generating high-fidelity egocentric videos with realistic hand-object interactions.
Chenyangguang Zhang, Botao Ye, Boqi Chen +4
1ETH Zurich · 2ETH AI Center · 3MPI for Informatics +1
Foundation video generation models such as WAN 2.2 exhibit strong text- and image-conditioned synthesis abilities but remain constrained to the same-view generation setting. In this work, we introduce Exo2EgoSyn, an adaptation of WAN 2.2 that unlocks Exocentric-to-Egocentric(Exo2Ego) cross-view video synthesis. Our framework consists of three key modules. Ego-Exo View Alignment(EgoExo-Align) enforces latent-space alignment between exocentric and egocentric first-frame representations, reorienting the generative space from the given exo view toward the ego view. Multi-view Exocentric Video Conditioning (MultiExoCon) aggregates multi-view exocentric videos into a unified conditioning signal, extending WAN2.2 beyond its vanilla single-image or text conditioning. Furthermore, Pose-Aware Latent Injection (PoseInj) injects relative exo-to-ego camera pose information into the latent state, guiding geometry-aware synthesis across viewpoints. Together, these modules enable high-fidelity ego view video generation from third-person observations without retraining from scratch. Experiments on ExoEgo4D validate that Exo2EgoSyn significantly improves Ego2Exo synthesis, paving the way for scalable cross-view video generation with foundation models. Source code and models will be released publicly.