InfiniHand: Streaming World-Space Hand Motion Estimation from Egocentric Video
Authors: Kerui Ren, Kaiwen Song, Weiguang Zhao, Yuxi Wang, Yufei Liu, Bo Dai, Haoyu Guo, Chunhua Shen, +3 more
Organizations: Shanghai Artificial Intelligence Laboratory · Shanghai Jiao Tong University · University of Science and Technology of China · University of Liverpool · Nanyang Technological University · The University of Hong Kong · Zhejiang University
World-space hand motion estimation from egocentric video requires recovering 3D articulated hand geometry while tracking camera egomotion. Existing approaches heavily rely on cascading independent hand pose estimators and SLAM systems, resulting in error accumulation, complex pipelines, and severe computational overhead. To address these limitations, we present InfiniHand, an end-to-end streaming feed-forward framework that jointly estimates MANO parameters, camera trajectories, and hand locations directly from uncalibrated egocentric video. InfiniHand integrates persistent spatiotemporal memory with hand-centered visual features, explicitly coupling camera motion with local hand geometry within a unified architecture. We train InfiniHand in two progressive stages by first learning robust camera-space hand priors and then extending to streaming world-space reconstruction. To support this process, we aggregate a pretraining corpus of approximately 5,000 hours of egocentric data across multiple public datasets. Extensive evaluations demonstrate that InfiniHand outperforms state-of-the-art baselines on in-domain benchmarks, achieving a 21.4% reduction in ARCTIC PA-p compared to ViDiHand while substantially mitigating world-space drift. Furthermore, InfiniHand generalizes robustly to in-the-wild videos and operates at 11.19 FPS, delivering more than twice the throughput of HaWoR.
Figures & tables
Figure 1: InfiniHand is a streaming feed-forward framework for accurate and efficient world-space hand estimation, pretrained on approximately 5,000 hours of egocentric video. Project page: https://infinihand.github.io/ .
Figure 2: Overview of InfiniHand. Stage I learns hand localization and MANO reconstruction from geometric and appearance features, transforming hand-frame predictions into camera coordinates. Stage II jointly estimates hands and camera trajectories with streaming memory, followed by sparse bundle adjustment for camera refinement and world-space reconstruction.
Method
ARCTIC
HOT3D
FAcc ↑
Recall ↑
F1 ↑
MP-p ↓
PA-p ↓
EPE-p ↓
GO-p ↓
CT-p ↓
FAcc ↑
Recall ↑
F1 ↑
MP-p ↓
PA-p ↓
EPE-p ↓
GO-p ↓
CT-p ↓
InterWild
0.878
0.943
0.959
30.82
15.95
53.89
25.39
0.097
0.669
0.881
0.868
77.17
24.81
71.48
58.50
0.213
HaMeR
0.875
0.943
0.957
29.20
14.60
65.29
24.91
0.095
0.692
0.904
0.883
68.31
21.46
59.08
49.64
0.102
Hamba
0.833
0.912
0.941
31.23
17.17
87.05
27.82
0.110
0.632
0.828
0.853
71.73
29.62
107.63
56.53
0.128
WildHands
0.879
0.946
0.960
25.70
13.94
50.52
22.32
0.058
0.655
0.863
0.844
52.79
28.95
111.44
53.93
0.157
OmniHands
0.866
0.949
0.954
29.67
14.20
51.51
24.58
0.087
0.649
0.895
0.868
63.28
22.68
68.44
49.12
0.133
Table 1: Camera-space quantitative comparison. Evaluating detection and camera-space hand motion metrics across four benchmarks, InfiniHand consistently yields the lowest PA-p. Bold and underlined denote best and second-best results.
Method
ARCTIC
HOT3D
EgoDex
PA-MPJPE ↓
W-MPJPE ↓
WA-MPJPE ↓
PA-MPJPE ↓
W-MPJPE ↓
WA-MPJPE ↓
PA-MPJPE ↓
W-MPJPE ↓
WA-MPJPE ↓
WiLoR-SLAM
7.23
65.86
46.31
6.46
106.77
45.20
10.25
96.02
40.95
HaWoR
9.03
95.39
45.45
5.86
93.70
35.02
10.36
103.85
35.31
Dyn-HaMR
10.85
114.10
63.60
10.17
296.72
132.08
11.58
78.41
37.11
Ours
7.07
59.21
43.59
5.71
87.37
35.76
4.91
28.77
16.47
Table 2: World-space quantitative comparison. We evaluate world-space hand motion and trajectory metrics across three benchmarks, with InfiniHand achieving the lowest W-MPJPE.
Figure 3: Camera-space qualitative comparison. InfiniHand recovers both hands without missed detections while producing precise hand motion under severe occlusion and in-the-wild scenes.
Figure 4: World-space qualitative comparison. Visual results on diverse datasets show that our reconstructed articulation and trajectories achieve the highest fidelity to ground truth.
Variant
Detection
Camera-space
World-space
FAcc ↑
Recall ↑
F1 ↑
MP-p ↓
PA-p ↓
EPE-p ↓
GO-p ↓
CT-p ↓
PA-MPJPE ↓
W-MPJPE ↓
WA-MPJPE ↓
w/o WiLoR Features
0.993
0.996
0.998
33.31
10.91
32.89
22.08
0.087
7.46
60.36
45.12
w/o LSP
0.993
0.996
0.998
18.63
8.02
26.30
12.97
0.064
7.21
59.57
44.35
w/o BA
0.993
0.996
0.998
17.09
7.72
23.13
12.80
0.044
7.07
80.48
51.20
w/o Stage II
0.985
0.989
0.990
17.29
7.81
23.56
13.01
0.045
7.40
185.76
111.05
w/o Mask Head
0.700
0.817
0.895
33.31
25.50
167.66
36.97
0.143
7.43
60.39
45.26
Table 3: Quantitative ablation study on ARCTIC. Results show the contributions of appearance features, translation recovery, sparse BA, joint training, and hand localization.
Appendix figures & tables2 assets
Supplementary material from the paper’s appendix.
Appendix
Figure 5: Additional in-the-wild camera-space results. Comparisons on Ego4D and Xperience illustrate hand reconstruction under diverse viewpoints, occlusions, and interactions.
Figure 6: Additional world-space qualitative results. Comparisons on ARCTIC, HOT3D, and EgoDex show reconstructed hand configurations and motion across further sequences.
Recovering camera and hand motion in world coordinates from egocentric video is a key capability for activity understanding, robot learning, and augmented reality. Existing systems typically decompose this problem into separate stages for camera motion, depth estimation, hand reconstruction, and trajectory refinement, resulting in substantial computational overhead and preventing the joint modeling of camera and hand motion. We introduce MINT (Minting IN-the-Wild Trajectories), a foundation model for world-space hand motion reconstruction from ego-centric RGB video. From a single shared spatiotemporal video representation, MINT jointly predicts the camera trajectory, field of view (FoV), camera-frame hand states, and per-frame hand observability, and then produces world-space hand motion via explicit coordinate transformations. Training such a model at scale is challenging, since paired world-space camera and hand annotations are scarce. We therefore develop an open-source labeling EGOPIPELINE that converts large collections of public egocentric videos into structured camera-and-hand trajectory supervision. MINT is first pretrained on these large-scale pseudo-labels and then fine-tuned on a small set of high-quality camera-and-hand annotations. Across public benchmarks MINT approaches state-of-the-art accuracy without seeing either benchmark in training, reaching 0.945 frame accuracy, 13.646 mm PA-MPJPE-p and 55.058 px EPE-p for camera-frame bimanual reconstruction on HOT3D, 4.690 mm RPE-T and 0.284 degrees RPE-R for camera trajectory, and a 3.67x end-to-end speedup over the labeling pipeline that supervises it. We release the model, training and inference code, labeling pipeline, and a curated 1,021-hour egocentric trajectory dataset.
Zijie Zhu, Weiren Cai, Yizhou Wang +4
1ShanghaiTech University · 3Wuji Technology · 4The University of Hong Kong +2
Recovering world space 4D motion of two interacting hands from egocentric video is a fundamental capability for supervising robot policy learning, where wrist trajectories track the end-effector and finger articulations specify the grasp pose. Two major challenges arise in this setting: hands frequently leave the camera view for extended periods due to head motion, and persistent hand-object interactions cause severe occlusions of one or both hands. Existing methods uniformly condition on noisy hand motion observations without accounting for their per-frame reliability, leading to substantial performance degradation. Our key insight is that accurate world space hand motion estimation is tightly coupled with the quality of per-frame hand observations. To this end, we decompose the quality of hand motion observations extracted from an off-the-shelf hand pose estimator into four channels: wrist global translation and finger articulations for both hands. We propose StableHand, a quality-aware flow-matching framework conditioned on these four-channel quality signals, which are predicted by a learned quality network. We naturally incorporate the quality signals into the flow-matching process through a per-channel forward schedule, a quality-adjusted velocity target, AdaLN modulation of the DiT denoiser, and a quality-aware ODE initialization. This unified generative process preserves high-quality observations while reconstructing unreliable ones using a learned bimanual motion prior. Experiments on HOT3D and ARCTIC, two egocentric benchmarks featuring long missing-hand spans and persistent hand-object occlusions, show that StableHand achieves state-of-the-art performance across all reported metrics, reducing W-MPJPE by 20-25% compared to the strongest baseline, with the largest gains on heavily occluded ARCTIC sequences.
Huajian Zeng, Chaohua Yao, Yuantai Zhang +3
Mohamed bin Zayed University of Artificial Intelligence · University of Illinois at Urbana-Champaign · Imperial College London
Egocentric video has become a primary source of supervision for embodied models, and its value rests on recovering hand motion in world coordinates, which camera motion and hand occlusion make difficult. Existing reconstruction pipelines typically separate hand and scene estimation, leave interaction attributes to separate task-specific models, and invoke several models per video, so no prior reconstruction model estimates these attributes and throughput becomes a practical constraint on large-scale annotation. We therefore introduce EgoFound3R, a unified end-to-end model that estimates world-space hand geometry in a metric scale shared with the scene, and predicts point-wise interaction attributes, including visibility, contact, and distance. The model integrates three designs: (i) structured hand prompts that transfer pretrained geometric priors to world-space hand reconstruction; (ii) an explicit hand representation that decodes hand geometry and interaction attributes; and (iii) a shared-parameter multi-rate design that lowers inference cost. Together, these designs predict hand geometry and point-wise attributes in one pass. On OakInk-v2, TACO, and HOI4D, EgoFound3R reduces the mean per-joint position error (MPJPE) by 43.2%, 22.4%, and 11.6% over previous methods and predicts point-wise contact and distance alongside the geometry in the same pass, while attaining approximately 6x higher throughput.
Hongming Fu, Jingcheng Shi, Wenjia Wang +2
School of AI, Shanghai Jiao Tong University · Rutgers University · The University of Hong Kong +1