Egocentric video has become a primary source of supervision for embodied models, and its value rests on recovering hand motion in world coordinates, which camera motion and hand occlusion make difficult. Existing reconstruction pipelines typically separate hand and scene estimation, leave interaction attributes to separate task-specific models, and invoke several models per video, so no prior reconstruction model estimates these attributes and throughput becomes a practical constraint on large-scale annotation. We therefore introduce EgoFound3R, a unified end-to-end model that estimates world-space hand geometry in a metric scale shared with the scene, and predicts point-wise interaction attributes, including visibility, contact, and distance. The model integrates three designs: (i) structured hand prompts that transfer pretrained geometric priors to world-space hand reconstruction; (ii) an explicit hand representation that decodes hand geometry and interaction attributes; and (iii) a shared-parameter multi-rate design that lowers inference cost. Together, these designs predict hand geometry and point-wise attributes in one pass. On OakInk-v2, TACO, and HOI4D, EgoFound3R reduces the mean per-joint position error (MPJPE) by 43.2%, 22.4%, and 11.6% over previous methods and predicts point-wise contact and distance alongside the geometry in the same pass, while attaining approximately 6x higher throughput.
Figures & tables
Figure 1: EgoFound3R at a glance. End-to-end world-space hand reconstruction with point-wise interaction attributes; the radar panel summarizes Tab. 1 .
Figure 2: Pipeline vs. end-to-end.
Figure 3: Overview of EgoFound3R. Egocentric video input is encoded at the hand rate H , and the built-in DINOv3 features of the frozen VGGT- Ω backbone provide the shared visual features, from which an adapter builds structured root, joint, and surface prompts. A bidirectional temporal Transformer links the two hands across H , and pooling neighboring prompts yields global scene anchors G⊆H at an adjustable H:G ratio. The prompted VGGT- Ω aggregator drives the explicit hand head that decodes hand geometry, visibility, contact, and distance, and the scene and metric heads that place the hands in world coordinates.
Figure 4: Cross-rate prompting and point decoding. (a, b) Hand prompts run at rate H while global anchors G⊆H aggregate them at a configurable lower rate. (c, d) A shared point representation decodes geometry and point-wise attributes.
Figure 5: Root recovery, scale alignment, and outputs. (a) Root recovery from 2D rays, relative joints, and a learned depth prior. (b) Inference-time scale alignment, with camera-space hands and rotations fixed. (c) Outputs at requested times R ; (d) 195 → MANO densification; (e) scale-consistent clip stitching.
OakInk-v2 ⋅ TACO ⋅ HOI4D
Method
Raw ↓
RR ↓
PA ↓
Sim(3)↓
W ↓
WA ↓
WiLoR †
91.24
70.40
81.51
27.66
28.33
27.75
7.78
9.49
9.62
16.89
25.95
24.14
42.14
51.14
49.41
17.04
26.14
24.57
PAD-Hand †
107.16
72.97
75.35
32.23
30.30
29.77
9.63
11.19
10.12
21.37
29.21
24.79
49.01
57.06
51.92
21.62
29.60
25.04
EgoForce †
123.57
126.63
147.77
47.75
44.53
52.78
11.09
11.93
12.92
32.62
47.84
49.22
80.59
105.52
142.58
33.10
46.65
55.76
EgoFound3R †
–
–
–
–
27.49
36.68
63.01
17.89
24.06
23.30
Dyn-HaMR 100w
96.39
127.12
245.06
17.91
27.47
26.49
7.84
11.62
11.89
13.76
29.27
30.47
36.16
77.88
173.62
17.19
31.46
65.71
Table 1: Joint reconstruction (MPJPE, mm); each cell lists OakInk-v2, TACO, and HOI4D left to right, with † marking methods placed with ground-truth extrinsics and ‡ marking EgoFound3R with ground-truth intrinsics and predicted extrinsics. For EgoFound3R † the camera-space columns are identical to the default row and are left blank. 100w subsets are unranked, W/WA are ranked within each extrinsics group.
OakInk-v2
TACO
HOI4D
Method
P ↑
R ↑
F1 ↑
P ↑
R ↑
F1 ↑
P ↑
R ↑
F1 ↑
InteractVLM 100w
0.180
0.403
0.202
0.388
0.705
0.472
0.313
0.537
0.281
S 2 Contact
0.327
0.081
0.079
0.519
0.104
0.146
0.444
0.265
0.284
ContactOpt
0.334
0.196
0.154
0.557
0.197
0.259
0.518
0.477
0.435
EgoFound3R
0.510
0.665
0.493
0.654
0.777
0.697
0.612
0.835
0.675
Table 2: Joint-level contact prediction . Superscripts mark 100-window subsets. InteractVLM is shown for completeness and does not enter the ranking.
OakInk-v2
TACO
HOI4D
Method
P ↑
R ↑
F1 ↑
P ↑
R ↑
F1 ↑
P ↑
R ↑
F1 ↑
HVD
0.823
0.808
0.792
0.878
0.834
0.853
0.917
0.769
0.831
EgoFound3R
0.919
0.747
0.802
0.917
0.850
0.880
0.892
0.884
0.885
Table 3: Visibility prediction at 21 joints on OakInk-v2, TACO, and HOI4D. HVD predicts joint visibility only; vertex-level results are reported in Appendix G .
Figure 6: World-space qualitative comparison. Rows: two viewpoints of each clip; columns: the compared methods, rendered in the reconstructed world frame. Hands are pink (left) and blue (right), darker over time; † marks methods placed with ground-truth extrinsics.
Figure 7: Camera-space qualitative comparison. Top: hand geometry (left/right hands in pink/blue). Bottom: visibility, contact, and distance for EgoFound3R ( Pred ) and the ground truth ( GT ); green/red = visible/occluded, red/light blue = contact/no contact, and distance is a blue-to-red 0–50 mm map.
Variant
Raw ↓
RR ↓
W ↓
WA ↓
Contact F1 ↑
Vis. F1 ↑
MAE(All) ↓
MAE(Pred) ↓
MAE(GT) ↓
0.05B
109.71
36.74
81.37
42.66
0.542
0.800
41.47
45.30
49.54
0.1B
96.90
30.55
61.66
35.81
0.579
0.801
49.27
52.15
58.61
MANO
86.66
24.29
57.45
35.15
0.643
0.854
45.46
48.21
51.97
w/o Ω
82.59
23.49
53.38
32.33
0.661
0.856
36.19
40.42
42.52
1-way
78.82
23.70
54.16
32.14
0.628
0.861
32.79
35.52
37.29
S5
54.64
18.54
40.03
25.16
0.697
0.880
26.26
28.09
30.47
Table 4: Core design ablation on TACO; the F1 columns are measured at 21 joints, and S k is EgoFound3R at stride k .
Figure 8: Inference efficiency.
Appendix figures & tables26 assets
Supplementary material from the paper’s appendix.
Appendix
Figure 9: Detailed technical framework. Full token layout, geometric interaction, explicit decoding, and auxiliary branches.
linear warmup over 100 steps, then cosine decay to 10−5
Batch size
8 per device on 8× NVIDIA H20 (effective 64 )
Resolution
256×256
Training steps
3,200
Appendix
Table 5: Training configuration. Only the hand prompts, adapters, and prediction heads are optimized; the VGGT- Ω backbone stays frozen.
Method
G/H/R
H FPS ↑
R FPS ↑
VRAM (GiB) ↓
Total params. (B/M)
EgoFound3R configurations
EgoFound3R-A ( 5122 )
4/21/60
48.63
138.95
3.87
1.356B
EgoFound3R-B ( 5122 )
34/168/500
43.65
129.91
9.61
1.356B
EgoFound3R-C ( 5122 )
100/500/1498
28.69
85.95
22.57
1.356B
Contact methods: GT hand + object mesh + object 6DoF
ContactOpt
–
1396.03
–
2.34
1.42M
Appendix
Table 6: Inference efficiency on one NVIDIA H20. The G/H/R column lists the anchor, hand, and output rates; H FPS and R FPS are in frames per second, VRAM is in GiB, and parameter counts use B/M; timings follow the released configuration of each system.
Figure 10: Hand-rate efficiency across global strides. Only the global stride s∈{1,3,5} changes, with parameters, 512×512 input, and 2 s windows held fixed, so the panels isolate its effect on hand-rate throughput (H FPS), the anchor and output rates, forward latency, and peak memory.
Dataset
Train sequences
Eval clips
OakInk-v2
565
400
TACO
839
400
HOI4D
1,431
461
H2O
138
283
HOT3D Aria
136
400
ARCTIC
226
434
Appendix
Table 7: Dataset splits. Train counts source sequences in the training partition; eval counts the 2 s windows sampled from the test partition under the common protocol. STERA-10M contributes training data only.
21 joints: MPJPE
195 vertices: MPVPE
778 vertices: MPVPE
Method
Raw ↓
RR ↓
PA ↓
Sim(3)↓
W ↓
WA ↓
Raw ↓
RR ↓
PA ↓
Sim(3)↓
W ↓
WA ↓
Raw ↓
RR ↓
PA ↓
Sim(3)↓
W ↓
WA ↓
OakInk-v2
WiLoR †
91.24
27.66
7.78
16.89
42.14
17.04
91.16
26.00
7.55
16.90
41.84
17.06
91.46
26.78
7.44
16.82
42.02
16.97
PAD-Hand †
107.16
32.23
9.63
21.37
49.01
21.62
106.47
30.03
9.32
21.15
47.23
21.40
106.98
30.79
9.12
21.02
47.41
21.26
EgoForce †
123.57
47.75
11.09
32.62
80.59
33.10
123.63
45.07
10.95
33.19
81.00
33.65
123.95
45.82
10.64
32.93
81.19
33.40
EgoFound3R †
–
–
–
–
27.49
17.89
–
–
–
–
27.79
18.07
–
–
–
–
28.53
18.76
Appendix
Table 8: Hand geometry on OakInk-v2, TACO, and HOI4D in mm at 21 joints and 195/778 vertices; marks follow Table 1 .
21 joints
195 vertices
778 vertices
Method
MPJVE ↓
MPJAE ↓
MPMVE ↓
MPMAE ↓
MPVVE ↓
MPVAE ↓
OakInk-v2
WiLoR †
116.81
4.66
110.07
4.38
112.47
4.49
PAD-Hand †
124.78
4.29
118.80
4.08
120.99
4.16
EgoForce †
258.44
12.11
242.73
11.35
247.57
11.58
HaWoR
59.59
2.07
56.20
1.96
57.23
1.99
Appendix
Table 9: Temporal consistency on OakInk-v2, TACO, and HOI4D. Velocity (MPJVE/MPMVE/MPVVE) and acceleration (MPJAE/MPMAE/MPVAE) errors at 21 joints, 195 vertices, and 778 vertices are in mm/s and m/s 2 ; marks follow Table 1 .
21 joints
195 vertices
778 vertices
Method
MPJVE ↓
MPJAE ↓
MPMVE ↓
MPMAE ↓
MPVVE ↓
MPVAE ↓
OakInk-v2
Dyn-HaMR 100w
132.68
6.04
130.90
5.95
131.84
6.00
EgoFound3R 100w,Dyn
63.85
1.32
61.40
1.27
61.69
1.27
TACO
Dyn-HaMR 100w
339.05
13.55
334.51
13.36
336.74
13.47
Appendix
Table 10: Matched 100-window temporal consistency. Dyn-HaMR and EgoFound3R are evaluated on the same windows; errors are in mm/s and m/s 2 and marks follow Table 1 .
21 joints
195 vertices
778 vertices
Method
P ↑
R ↑
F1 ↑
P ↑
R ↑
F1 ↑
P ↑
R ↑
F1 ↑
OakInk-v2
InteractVLM 100w
0.180
0.403
0.202
0.173
0.408
0.205
0.168
0.412
0.201
S 2 Contact
0.327
0.081
0.079
0.401
0.180
0.158
0.397
0.198
0.173
ContactOpt
0.334
0.196
0.154
0.334
0.234
0.178
0.338
0.248
0.185
EgoFound3R
0.510
0.665
0.493
0.525
0.611
0.479
0.516
0.329
0.320
Appendix
Table 11: Contact prediction on OakInk-v2, TACO, and HOI4D at 21 joints, 195 vertices, and 778 vertices. Superscripts mark 100-window subsets, and marks follow Table 1 .
Dataset
Method
195 vertices
778 vertices
P ↑
R ↑
F1 ↑
P ↑
R ↑
F1 ↑
OakInk-v2
EgoFound3R
0.556
0.705
0.595
0.555
0.690
0.588
TACO
EgoFound3R
0.437
0.818
0.567
0.439
0.805
0.565
HOI4D
EgoFound3R
0.572
0.789
0.660
0.581
0.781
0.663
Appendix
Table 12: Vertex-level visibility at the native 195 vertices and the interpolated 778-vertex mesh.
Dataset
Method
AP ↑
best-F1 ↑
F1@0.5 ↑
OakInk-v2
ContactOpt
0.308
0.411 (0.21)
0.178
S 2 Contact
0.395
0.449 (0.09)
0.158
EgoFound3R
0.603
0.617 (0.39)
0.479
TACO
ContactOpt
0.475
0.542 (0.15)
0.311
S 2 Contact
0.512
0.558 (0.09)
0.285
EgoFound3R
0.714
0.691 (0.43)
0.661
Appendix
Table 13: Decision-threshold sweep for contact prediction on the 195 vertices. AP and best-F1 pool the points of a dataset over 99 thresholds, with the best-F1 threshold in parentheses; marks follow Table 1 .
our GT
2 mm band
Dataset
Method
P ↑
R ↑
F1 ↑
P ↑
R ↑
F1 ↑
OakInk-v2
ContactOpt
0.334
0.234
0.178
0.080
0.335
0.096
S 2 Contact
0.401
0.180
0.158
0.117
0.331
0.115
EgoFound3R
0.525
0.611
0.479
0.103
0.741
0.164
TACO
ContactOpt
0.571
0.249
0.311
0.101
0.358
0.140
S 2 Contact
0.590
0.210
0.285
0.127
0.376
0.171
Appendix
Table 14: Narrow-band contact comparison on the 195 vertices at the 0.5 operating point. The two column groups differ only in the labeling tolerance; the 2 mm band is defined on the 195 vertices only.
Figure 11: Where each method’s predicted contacts lie. Predicted-positive distances on the 195 vertices, one row per dataset. Left: cumulative fraction of a method’s positives within a given distance to the surface, ground truth in gray, with vertical lines marking the 2 mm coverage band of the two geometric baselines and our 14 mm contact band. Right: recall of ground-truth contacts as the labelling threshold varies, with reference lines at 2, 14, and 18 mm.
21 joints: MPJPE
195 vertices: MPVPE
778 vertices: MPVPE
Method
Raw ↓
RR ↓
PA ↓
Sim(3)↓
W ↓
WA ↓
Raw ↓
RR ↓
PA ↓
Sim(3)↓
W ↓
WA ↓
Raw ↓
RR ↓
PA ↓
Sim(3)↓
W ↓
WA ↓
OakInk-v2
Dyn-HaMR 100w
96.39
17.91
7.84
13.76
36.16
17.19
95.78
16.87
7.67
13.73
35.84
17.20
95.94
17.22
7.56
13.62
35.90
17.06
EgoFound3R 100w,Dyn
50.88
21.17
8.50
20.30
30.20
20.88
49.07
20.97
8.54
20.49
30.21
21.02
49.17
22.25
9.53
21.11
30.99
21.63
EgoFound3R ‡,100w,Dyn
43.29
21.18
8.51
21.15
32.68
21.88
42.73
20.97
8.54
21.33
32.71
22.06
42.94
22.28
9.55
21.93
33.40
22.63
TACO
Appendix
Table 15: Matched 100-window hand geometry. Dyn-HaMR and EgoFound3R are evaluated on the same windows; errors are in mm.
Method
21 joints
195 vertices
778 vertices
P ↑
R ↑
F1 ↑
P ↑
R ↑
F1 ↑
P ↑
R ↑
F1 ↑
OakInk-v2
InteractVLM 100w
0.180
0.403
0.202
0.173
0.408
0.205
0.168
0.412
0.201
EgoFound3R 100w,InteractVLM
0.381
0.762
0.431
0.441
0.638
0.442
0.457
0.295
0.298
TACO
InteractVLM 100w
0.388
0.705
0.472
0.348
0.743
0.451
0.356
0.750
0.458
Appendix
Table 16: Matched 100-window contact comparison. InteractVLM and EgoFound3R are evaluated on exactly the same windows, and the best value in each column is marked in bold.
21 joints: MPJPE
195 vertices: MPVPE
778 vertices: MPVPE
Variant
Raw ↓
RR ↓
PA ↓
Sim(3)↓
W ↓
WA ↓
Raw ↓
RR ↓
PA ↓
Sim(3)↓
W ↓
WA ↓
Raw ↓
RR ↓
PA ↓
Sim(3)↓
W ↓
WA ↓
OakInk-v2
0.05B
112.47
35.90
16.76
33.35
65.09
35.51
110.83
35.50
16.15
33.45
65.83
35.66
111.13
36.33
16.20
33.46
66.14
35.63
0.1B
98.34
31.74
15.24
30.74
61.31
32.39
98.33
30.92
15.27
30.67
59.29
32.32
98.69
31.64
15.36
30.73
59.56
32.32
MANO
73.39
21.02
8.74
24.13
45.73
25.71
73.27
19.87
8.55
24.30
45.46
25.92
73.12
21.01
9.36
24.70
46.05
26.27
w/o Ω
86.13
22.06
9.92
23.63
44.03
24.88
85.25
20.59
9.54
23.69
43.87
24.95
86.07
21.67
10.32
24.17
44.39
25.40
Appendix
Table 17: Hand-geometry ablation on OakInk-v2, TACO, and HOI4D. Errors are in mm; S k is EgoFound3R at stride k .
21 joints
195 vertices
778 vertices
Variant
P ↑
R ↑
F1 ↑
VP ↑
VR ↑
VF1 ↑
P ↑
R ↑
F1 ↑
VP ↑
VR ↑
VF1 ↑
P ↑
R ↑
F1 ↑
VP ↑
VR ↑
VF1 ↑
OakInk-v2
0.05B
0.541
0.283
0.298
0.820
0.736
0.760
0.600
0.251
0.277
0.678
0.170
0.244
0.373
0.165
0.175
0.695
0.161
0.235
0.1B
0.549
0.358
0.345
0.843
0.708
0.752
0.620
0.170
0.205
0.720
0.192
0.272
0.325
0.194
0.185
0.733
0.186
0.267
MANO
0.512
0.595
0.474
0.927
0.807
0.856
0.537
0.582
0.485
0.566
0.775
0.642
0.530
0.175
0.206
0.567
0.760
0.636
w/o Ω
0.535
0.606
0.480
0.925
0.796
0.844
0.552
0.557
0.463
0.554
0.770
0.630
0.276
0.250
0.203
0.553
0.757
0.625
Appendix
Table 18: Contact and visibility ablation on OakInk-v2, TACO, and HOI4D. P, R, and F1 denote contact precision, recall, and F1, and VP, VR, and VF1 the corresponding visibility scores.
21 joints
195 vertices
778 vertices
Variant
MAE(All) ↓
MAE(Pred) ↓
MAE(GT) ↓
MAE(All) ↓
MAE(Pred) ↓
MAE(GT) ↓
MAE(All) ↓
MAE(Pred) ↓
MAE(GT) ↓
OakInk-v2
0.05B
45.01
47.86
55.45
44.29
46.07
55.38
43.88
25.49
54.74
0.1B
39.91
41.60
45.88
39.61
41.08
47.73
39.35
25.20
46.76
MANO
34.73
36.28
39.00
35.32
37.76
41.37
34.66
17.35
41.13
w/o Ω
32.79
30.73
31.78
32.24
31.80
33.21
32.28
28.37
32.47
Appendix
Table 19: Distance MAE ablation on OakInk-v2, TACO, and HOI4D at 21 joints, 195 vertices, and 778 vertices. Errors are in mm; MAE(All/Pred/GT) use all valid, predicted-contact, and ground-truth-contact points.
Method
ATE (m) ↓
Rot ( ∘ ) ↓
AbsRel ↓
RMSE (m) ↓
δ1↑
sd
HaWoR
0.0064
1.22
–
–
–
–
Dyn-HaMR 100w
0.0064
1.26
–
–
–
–
ReViV
0.0068
1.65
0.2649
0.3577
0.5829
0.6401
EgoFound3R
0.0055
0.99
0.1318
0.2015
0.8469
1.1364
Appendix
Table 20: Scene estimation on TACO for hand-centric world-space methods. Depth follows the scale-aligned protocol defined above; 100w rows are unranked.
Variant
ATE (m) ↓
Rot ( ∘ ) ↓
AbsRel ↓
RMSE (m) ↓
δ1↑
0.05B
0.0086
1.86
0.117
0.232
0.857
0.1B
0.0074
1.83
0.122
0.256
0.825
MANO
0.0084
1.60
0.132
0.256
0.806
w/o Ω
0.0069
1.15
0.178
0.242
0.728
1-way
0.0069
1.15
0.178
0.242
0.728
S1
0.0043
0.82
0.136
0.284
0.805
Appendix
Table 21: Scene ablation on TACO.
Figure 12: Qualitative comparison on ARCTIC. Top: hand geometry. Bottom: per-point visibility, contact, and distance for EgoFound3R ( Pred ) and the ground truth ( GT ).
Figure 13: Qualitative comparison on OakInk-v2. Top: hand geometry. Bottom: per-point visibility, contact, and distance for EgoFound3R ( Pred ) and the ground truth ( GT ).
Figure 14: Qualitative hand geometry on TACO. Columns as in Fig. 12 ; the matching attribute panel of the same windows is shown separately below.
Figure 15: Per-point attributes on TACO. Rows follow the same windows and order as Fig. 14 ; columns show the RGB input and the per-point visibility, contact, and distance for EgoFound3R ( Pred ) and the ground truth ( GT ).
Figure 16: World-space qualitative comparison, part 1.
Figure 17: World-space qualitative comparison, part 2.
Egocentric video offers scalable manipulation data for embodied AI, yet recovering metric 3D hand trajectories remains challenging due to severe object occlusion and frequent out-of-sight gaps. Existing single-frame and windowed temporal regressors fail when a hand shortly leaves the frame, while recent video diffusion models (VDMs) rely on heavy, stochastic multi-step sampling as pixel-space renderers. We instead repurpose VDM into a deterministic geometry encoder. A single forward pass over the clean latent exposes scene content beyond current observations, including occluded and out-of-sight hands. We introduce ACE-Ego-Hand, an offline clip-level framework that extracts features via a Deterministic Clean-Latent Encoder and decodes them with a Bidirectional Spatiotemporal Decoder. ACE-Ego-Hand recovers continuous bimanual trajectories with metric placement and no external detector, while a Ray-Based Camera Solver supports a second configuration that requires no test-time camera intrinsics. Across five egocentric benchmarks, ACE-Ego-Hand sets a new state of the art, cutting MPJPE-p by 30% on occlusion-heavy ARCTIC and 40% on HOT3D. These gains reach 46%-61% once out-of-sight hands are included in the evaluation, offering a scalable path from everyday human video to robot manipulation data.
Yufei Liu, Xixi Wang, Hao Li +8
1Shanghai Jiao Tong University · 2Nanyang Technological University · 3The Chinese University of Hong Kong +1
Human dexterity is guided by two eyes watching two hands: binocular vision supplies the metric 3D structure that fine-grained manipulation consumes. Egocentric stereo is therefore the natural perceptual interface for robots, AR, and VR-yet metric 3D hand reconstruction from this very signal still has neither an end-to-end model nor an in-the-wild benchmark. We propose ESTHER, a model whose stereo geometry, temporal reasoning, and output representation are designed for wearable egocentric stereo. It is trained on pseudo-labels from a calibrated labeling pipeline and in turn assembles our benchmark ESTHER3D, an egocentric stereo hand dataset pairing a large in-the-wild training set of model-generated labels with a motion capture test set of true metric ground truth. Experiments show state-of-the-art accu?racy, superior external generalization, and robustness to the missing views, dropped frames, and lighting and motion blur extremes of real egocentric capture that break existing meth?ods. This robustness runs deeper than graceful degradation: stereo guidance teaches the model to bind apparent hand scale to metric depth, so it not only adapts to different stereo rigs and modalities with minimal fine-tuning, but more strikingly preserves true metric scale even after collapsing to a single monocular view.
World-space hand motion estimation from egocentric video requires recovering 3D articulated hand geometry while tracking camera egomotion. Existing approaches heavily rely on cascading independent hand pose estimators and SLAM systems, resulting in error accumulation, complex pipelines, and severe computational overhead. To address these limitations, we present InfiniHand, an end-to-end streaming feed-forward framework that jointly estimates MANO parameters, camera trajectories, and hand locations directly from uncalibrated egocentric video. InfiniHand integrates persistent spatiotemporal memory with hand-centered visual features, explicitly coupling camera motion with local hand geometry within a unified architecture. We train InfiniHand in two progressive stages by first learning robust camera-space hand priors and then extending to streaming world-space reconstruction. To support this process, we aggregate a pretraining corpus of approximately 5,000 hours of egocentric data across multiple public datasets. Extensive evaluations demonstrate that InfiniHand outperforms state-of-the-art baselines on in-domain benchmarks, achieving a 21.4% reduction in ARCTIC PA-p compared to ViDiHand while substantially mitigating world-space drift. Furthermore, InfiniHand generalizes robustly to in-the-wild videos and operates at 11.19 FPS, delivering more than twice the throughput of HaWoR.
Kerui Ren, Kaiwen Song, Weiguang Zhao +8
Shanghai Artificial Intelligence Laboratory · Shanghai Jiao Tong University · University of Science and Technology of China +4