Recovering faithful 3D hand motion from video remains challenging due to frequent occlusions and incomplete visual observations, which make frame-wise pose estimates unreliable and temporally inconsistent. To address this problem, we propose JoHan, a unified generative framework that recovers hand motion directly from video sequences without relying on intermediate per-frame pose predictions. Trained from scratch, our model jointly generates aligned 2D and 3D local hand pose sequences by learning their temporal dynamics and cross-representation correspondence. The generated 2D trajectories exploit direct spatial and temporal cues from the 2D images to guide the following generative 3D motion reconstruction, while the learned motion prior promotes temporal consistency. Their learned 2D-3D correspondence further enables recovery of the hand's global position and orientation relative to the camera. Extensive experiments on challenging benchmarks demonstrate significantly improved accuracy and speed in local hand-pose and camera-space reconstruction. Notably, our method captures much better hand-motion dynamics, producing significantly smoother motion than previous methods while maintaining high per-frame pose accuracy.
Figures & tables
Figure 1: Joint motion generation and qualitative comparisons. Left: prior methods typically use per-frame pose estimates to condition or initialize global motion sequence reconstruction. JoHan jointly generates local 2D–3D motion sequences from video, followed by the global reconstruction. Right: world-space mesh sequences recovered from two videos from HO3D with a static camera (top, compared with Dyn-HaMR ( Yu et al., 2025 ) ) and two from HOT3D with a moving camera (bottom, compared with HaWoR ( Zhang et al., 2025 ) ). Methods are denoted by color: Ground truth , JoHan , Dyn-HaMR , and HaWoR . Selected input frames appear in temporal order from top to bottom.
Figure 2: Overview of JoHan. We first combine full-frame video features with encoded detector-derived hand masks m to form visual conditions F . Within the Joint 2D–3D Motion Reconstruction module, the 2D generator predicts clean local motion X^s2D from the noisy state Xs2D at each flow step. This prediction is then encoded as guidance Gs for the 3D generator, allowing the predicted 2D geometry to inform 3D motion generation. After that, detector-derived bounds restore the image-plane location and scale of the generated 2D local poses. A PnP solver then uses these restored poses, the corresponding local MANO joints, and camera intrinsics K to recover global orientation R^ and translation Γ^ , which are used to place the reconstructed hands in the camera coordinate system.
Method
PA-MPJPE ↓
PA-MPVPE ↓
AUC J ↑
F@5 ↑
F@15 ↑
Err2D↓
Deformer
9.4
9.1
–
0.546
0.963
–
HandOccNet
9.1
8.8
0.819
0.564
0.963
–
AMVUR
8.3
8.2
0.835
0.608
0.965
–
HaMeR
7.7
7.9
0.846
0.635
0.980
6.51
WiLoR
7.5
7.7
0.851
0.646
0.983
6.37
JoHan
7.4
7.5
0.851
0.648
0.983
3.92
Table 1: Hand reconstruction on HO3D. Published v2 baseline 3D scores; v3 HaMeR/WiLoR re-evaluated by us. 3D errors are in mm; Err2D is 2D MPJPE in px at 256×256 . Dashes: unavailable. Bold / underlined : best/second-best before rounding in each HO3D subtable and Table 3 .
Table 4
Variant
PA-MPJPE (mm) ↓
Err2D (px) ↓
Full model
7.8
5.21
Auxiliary supervision
Without 3D joint loss
7.9
5.32
Without reprojection loss
8.2
5.96
Use of detections
Cropped video
8.9
15.68
Table 4: HO3Dv3 ablations. All variants use DINOv3 ViT-B/16; the full model is repeated as the reference for both groups.
Appendix figures & tables8 assets
Supplementary material from the paper’s appendix.
Appendix
Figure 3: Qualitative hand reconstruction from videos in HO3D. Each group shows one selected frame. Each method’s row contains the input RGB frame, its predicted mesh projected onto the hand crop, and two camera-space mesh views overlaid with ground truth. Ground truth is green; JoHan, HaMeR ( Pavlakos et al., 2024 ) , and WiLoR ( Potamias et al., 2025 ) are blue, orange, and purple, respectively.
Figure 4: Qualitative hand reconstruction from videos in HOT3D. We show six selected frames with hand–object occlusion. Each row shows the input RGB frame, a mesh projection onto the hand crop, and two camera-space mesh views overlaid with ground truth. The three method rows compare JoHan, HaMeR ( Pavlakos et al., 2024 ) , and WiLoR ( Potamias et al., 2025 ) . The layout and colors follow Figure 3 .
Figure 5: World-space motion comparison on HOT3D. Four examples from two sequences compare ground truth (gray), JoHan (blue), and HaWoR ( Zhang et al., 2025 ) (red). The enlarged views project initially aligned world-space meshes and trajectories onto the first two ground-truth principal axes. Grid spacing is 10 cm. Each row shows two frames from one sequence, indexed from zero. The curves show per-frame W-MPJPE; vertical markers indicate the displayed frames.
Figure 6: Conditioned generator architecture. (a) Each generator maps noised local motion Xsd to a clean-motion prediction X^sd . Only the 3D generator uses 2D-Guided Fusion, which incorporates the guidance Gs from the current 2D prediction before motion tokenization. (b) Each spatiotemporal block applies visual cross-attention conditioned on F , followed by spatial and temporal self-attention. Gold dashed arrows indicate AdaLN-Zero conditioning from timestep s and hand side h in both generators.
Component
2D generator
3D generator
Motion tokens per frame
21
15 pose + 1 shape
Token dimension
384
384
Spatiotemporal blocks
4
6
Attention heads
4
4
Feedforward hidden dimension
384
1536
2D-Guided Fusion
–
Gs at each flow step
Appendix
Table 5: Motion generator configurations. Both generators use the same full-frame visual conditions and timestep/hand-side conditioning, with independent attention parameters.
Figure 11
Bbox refiner
PA-MPJPE ↓
MPJPE ↓
W-MPJPE ↓
WA-MPJPE ↓
AccEr ↓
Err2D↓
Without
5.20
46.05
78.37
26.86
3.42
5.18
With
4.75
33.84
55.05
20.36
2.69
3.27
Appendix
Table 6: Bounding-box refiner ablation on HOT3D. Both settings are evaluated on the same 42,990 frames. Joint errors are in mm, AccEr is in m/s2 , and Err2D is in pixels. Lower values are better for all metrics.
Method
PA-MPJPE ↓
W-MPJPE ↓
WA-MPJPE ↓
AccEr ↓
HaMeR
9.00
145.15
36.31
11.07
WiLoR
6.28
72.81
24.72
9.48
HaWoR
5.47
42.95
13.43
5.13
JoHan
4.70
45.78
14.89
3.34
Appendix
Table 7: HOT3D results with ground-truth bounding boxes. All rows report our evaluations on 100-frame segments using ground-truth hand boxes. Joint errors are in mm and AccEr is in m/s2 . JoHan uses HaWoR’s camera motion estimation module. Lower values are better for all metrics.
Accurate monocular 4D hand reconstruction remains challenging. Per-frame discriminative regressors lack temporal context and often produce jittery predictions. Temporal models improve consistency by aggregating information across frames, but they are typically deterministic regressors, making them vulnerable to ambiguous observations caused by occlusion and motion blur. Generative modeling offers a natural alternative by learning a prior over plausible hand motion sequences, enabling coherent hand-state recovery when visual evidence is incomplete or unreliable. Motivated by this observation, we present HandFlow, a fully generative flow-matching framework for temporally coherent 3D hand pose and shape estimation from monocular video. Given visual and skeletal observations, HandFlow denoises an entire temporal window of MANO parameters through a single ODE integration. To support this, we use a Flux-style dual-stream transformer that attends across the full sequence to capture long-range dependencies without autoregressive decoding, and a confidence-aware continuous masking mechanism that blends observed features with learnable mask tokens to handle noisy or missing observations. Experiments on DexYCB and HOT3D show that HandFlow achieves state-of-the-art performance, with particularly large gains in world-space accuracy and temporal smoothness. It reduces world-space pose error by over 30% compared with the strongest baseline and achieves the lowest acceleration error among all evaluated methods, while remaining competitive in per-frame pose accuracy. Moreover, on a single GPU HandFlow reconstructs a 150-frame sequence at 47 fps, about 12x faster than the fastest prior video-based method, with reconstruction itself accounting for only a small fraction of the end-to-end latency.
Mingxi Xu, Bowen Duan, Yi Gu +3
The Hong Kong University of Science and Technology (Guangzhou), China · Google, USA
Egocentric video offers scalable manipulation data for embodied AI, yet recovering metric 3D hand trajectories remains challenging due to severe object occlusion and frequent out-of-sight gaps. Existing single-frame and windowed temporal regressors fail when a hand shortly leaves the frame, while recent video diffusion models (VDMs) rely on heavy, stochastic multi-step sampling as pixel-space renderers. We instead repurpose VDM into a deterministic geometry encoder. A single forward pass over the clean latent exposes scene content beyond current observations, including occluded and out-of-sight hands. We introduce ACE-Ego-Hand, an offline clip-level framework that extracts features via a Deterministic Clean-Latent Encoder and decodes them with a Bidirectional Spatiotemporal Decoder. ACE-Ego-Hand recovers continuous bimanual trajectories with metric placement and no external detector, while a Ray-Based Camera Solver supports a second configuration that requires no test-time camera intrinsics. Across five egocentric benchmarks, ACE-Ego-Hand sets a new state of the art, cutting MPJPE-p by 30% on occlusion-heavy ARCTIC and 40% on HOT3D. These gains reach 46%-61% once out-of-sight hands are included in the evaluation, offering a scalable path from everyday human video to robot manipulation data.
Yufei Liu, Xixi Wang, Hao Li +8
1Shanghai Jiao Tong University · 2Nanyang Technological University · 3The Chinese University of Hong Kong +1
4D hand motion reconstruction from egocentric video is bottlenecked by clear limitations of existing methods: image-based pipelines depend on a detector that fails under heavy occlusion, while video-based methods rely on temporal modules learned only from scarce hand-pose annotations, a narrow signal insufficient to model motion dynamics, occlusion reasoning, and hand-object interaction. These capabilities, however, are exactly what video generative models must implicitly acquire when trained to synthesize coherent video at internet scale. Motivated by this, we present ViDiHand, which leverages the representations of a pretrained video diffusion model to reconstruct 4D two-hand pose. We adapt it via a hand-overlay rendering objective that specializes its features for hands while preserving its world priors. A decoder then recovers metric-scale pose from the adapted features. The whole pipeline operates directly on full frames--no detector, no infiller, and no test-time optimization. On ARCTIC, HOT3D, and HOI4D, ViDiHand substantially outperforms prior methods, establishing video diffusion models as a powerful new foundation for hand motion reconstruction and a promising route to scalable in-the-wild data collection for embodied AI. Project page: https://vidihand.github.io.
Yuxi Wang, Chengkai Jin, Yufei Liu +6
1Nanyang Technological University · 2Shanghai Jiao Tong University