cs.CVOct 7, 2026

Video-Conditioned Generative Joint 2D-3D Hand Motion Recovery

Authors: Chen Xu, Yunqi Li, Binbin Huang, Brent Yi, Shenghua Gao, Yi Ma

Organizations: The University of Hong Kong · UC Berkeley

Abstract

Recovering faithful 3D hand motion from video remains challenging due to frequent occlusions and incomplete visual observations, which make frame-wise pose estimates unreliable and temporally inconsistent. To address this problem, we propose JoHan, a unified generative framework that recovers hand motion directly from video sequences without relying on intermediate per-frame pose predictions. Trained from scratch, our model jointly generates aligned 2D and 3D local hand pose sequences by learning their temporal dynamics and cross-representation correspondence. The generated 2D trajectories exploit direct spatial and temporal cues from the 2D images to guide the following generative 3D motion reconstruction, while the learned motion prior promotes temporal consistency. Their learned 2D-3D correspondence further enables recovery of the hand's global position and orientation relative to the camera. Extensive experiments on challenging benchmarks demonstrate significantly improved accuracy and speed in local hand-pose and camera-space reconstruction. Notably, our method captures much better hand-motion dynamics, producing significantly smoother motion than previous methods while maintaining high per-frame pose accuracy.

Figures & tables

Appendix figures & tables8 assets

Supplementary material from the paper’s appendix.

Appendix

Explore similar work

CardsList
  1. HandFlow: Fully Generative 4D Hand Recovery with Flow Matching

    Jul 13, 2026Mingxi Xu, Bowen Duan, Yi Gu +3Articulated Object Reconstruction4D Reconstruction

  2. ACE-Ego-Hand: Repurposing Video Diffusion Models for Occlusion-Robust Egocentric 3D Hand Motion Recovery

    Aug 20, 2026Yufei Liu, Xixi Wang, Hao Li +8Video Diffusion Models3D Hand Pose Estimation

  3. The Surprising Effectiveness of Video Diffusion Models for Hand Motion Reconstruction

    Jun 29, 2026Yuxi Wang, Chengkai Jin, Yufei Liu +6Articulated Object ReconstructionVideo Diffusion Models