Exo2EgoHOI: Hand-Object-Interaction Aware Exocentric-to-Egocentric Video Generation
Organizations: MBZUAI · Zhejiang University · InSpatio
Abstract
Egocentric videos of human manipulation provide valuable visual experience for embodied intelligence, yet collecting such data at scale is costly. Exocentric-to-egocentric video generation offers a scalable alternative by transforming abundant third-person manipulation videos into first-person observations. However, existing methods often struggle to faithfully preserve demonstrated hand-object interactions (HOI) across large viewpoint changes due to insufficient fine-grained interaction guidance and weak object-centric anchoring. We present Exo2EgoHOI, an HOI-aware video generative framework for interaction-preserving exocentric-to-egocentric translation. To preserve fine-grained HOI, we introduce a unified 4D HOI prior that combines scene geometry, articulated hand renderings, and dense hand-object relation fields, together with a dual-branch residual adapter for injecting structural and relational cues into the video generation backbone. To preserve object consistency, we introduce Decomposed Gated Cross-Attention, which separately encodes object and background references and adaptively integrates global semantic and local appearance features as object-centric anchors. Experiments on ARCTIC-HOI and Ego-Exo4D demonstrate substantial improvements in object consistency and HOI preservation while maintaining competitive visual fidelity. In particular, on ARCTIC-HOI, Exo2EgoHOI improves object mIoU by 32.3% and reduces MPJPE and PA-MPJPE by 34.7% and 50.0%, respectively, relative to the respective best baseline results. Project page: https://rcl-robotics.github.io/Exo2EgoHOI/.
Figures & tables
| Visual Fidelity | Object Consistency | Hand Consistency | ||||||||
|---|---|---|---|---|---|---|---|---|---|---|
| Method | PSNR | SSIM | LPIPS | CLIP-I | mIoU | Center-Err. | CLIP-O | MPJPE | WA-MPJPE | PA-MPJPE |
| ARCTIC-HOI | ||||||||||
| WAN VACE ( Jiang et al., 2025 ) | 12.78 | 0.6163 | 0.5115 | 0.8485 | 0.2336 | 0.0893 | 0.8080 | 2870.22 | 371.18 | 20.36 |
| EgoWorld ( Park et al., 2026 ) | 11.90 | 0.4652 | 0.6203 | 0.7995 | 0.2241 | 0.1131 | 0.7869 | 2282.03 | 440.60 | 57.92 |
| Vista4D ( Lin et al., 2026b ) | 13.58 | 0.6614 | 0.4324 | 0.8737 | 0.2768 | 0.0929 | 0.8176 | 2123.78 | 387.73 | 28.91 |
| EgoX ( Kang et al., 2026 ) | 11.83 | 0.5941 | 0.4743 | 0.8910 | 0.2933 | 0.0804 | 0.8345 | 2019.81 | 325.56 | 22.12 |
| Methods | HOI Priors | DGCA | Evaluation Metrics | ||||||||||
|---|---|---|---|---|---|---|---|---|---|---|---|---|---|
| Scene | Hand | Inter. | Global | Local | Gate | LPIPS | CLIP-I | mIoU | CLIP-O | MPJPE | WA-MPJPE | PA-MPJPE | |
| #0 Baseline | ✗ | ✗ | ✗ | ✗ | ✗ | ✗ | 0.5577 | 0.8253 | 0.1148 | 0.8056 | 2631.66 | 371.25 | 41.42 |
| #1 w/o All Priors | ✗ | ✗ | ✗ | ✓ | ✓ | ✓ | 0.5351 | 0.8449 | 0.1378 | 0.8255 | 2254.26 | 346.98 | 21.11 |
| #2 w/ Scene | ✓ | ✗ | ✗ | ✓ | ✓ | ✓ | 0.4435 | 0.9008 | 0.3104 | 0.8391 | 1978.57 | 327.60 | 14.30 |
| #3 w/ Scene & Hand | ✓ | ✓ | ✗ | ✓ | ✓ | ✓ | 0.4063 | 0.9108 | 0.3323 | 0.8421 | 1570.51 | 315.28 | 13.33 |
| #4 w/o Decomp. CA | ✓ | ✓ | ✓ | ✗ | ✗ | ✗ | 0.3953 | 0.8853 | 0.3023 | 0.8198 | 2327.34 | 398.94 | 16.02 |
Appendix figures & tables5 assets
Supplementary material from the paper’s appendix.
Appendix
| Operation | Channels | Kernel | Stride | Output |
|---|---|---|---|---|
| Branch Conv 1–3 | ||||
| Branch Conv 4 | ||||
| Branch Conv 5 | ||||
| Branch Conv 6 | ||||
| Concatenate + fuse | ||||
| Patch projection |