Recently, approaches that leverage human video datasets for robot policy training have become increasingly prevalent. However, most existing hand trackers regress pose from cropped frames with limited priors on hand motion and object interaction, resulting in inaccurate and physically inconsistent estimates. Moreover, the lack of physical cues, e.g., contact and force, limits the use of human videos for robot policy training. To this end, we propose RLHND, a video foundation model-based hand tracking model that jointly estimates hand pose and realistic tactile information from monocular egocentric videos. RLHND turns the pre-trained Cosmos 3 video diffusion backbone into a deterministic clip-level feature extractor via clean-latent conditioning, carrying its learned priors on hand motion and hand-object interaction into tracking. For pose estimation, RLHND (i) predicts hand poses with anatomically plausible joint angles and (ii) enables optional conditioning on the shape parameter to maintain consistent hand shape within the same video and even across videos recorded by the same actor. For tactile estimation, a separate tactile expert stream, trained with the pose stream frozen, predicts dense contact and force over the hand surface. We further adopt LBS-based feature spreading to enable vertex-wise feature extraction without costly per-vertex attention. RLHND achieves state-of-the-art performance across various benchmark datasets for pose estimation, while also achieving state-of-the-art performance in contact and force estimation. Moreover, we demonstrate the utility of RLHND for robot learning through retargeting results and real-world robot experiments. The code will be publicly available at https://seungjun-moon.github.io/rlhnd/.
Figures & tables
Figure 1: 3D instability of hand motion reconstruction methods. Left: Wrist depth over 850 frames from an ARCTIC test-set clip. Baselines exhibit severe oscillation or drift relative to the ground-truth depth. Right: Per-video standard deviation of the estimated hand size. Even within a single video, existing hand motion estimators exhibit ∼ 8 mm variation in estimated hand size.
Figure 2: RLHND overview. A pre-trained Cosmos 3 backbone receives the clip through its conditioning-frame interface and is adapted with a trainable patch embedding and LoRA adapters. The pose expert stream, trained in stage-1, yields MANO parameters and a metric camera-space translation while the tactile expert stream is trained in stage-2 with the pose stream frozen.
Figure 3: Qualitative comparison of hand motion reconstruction. We visualize the reconstruction results from HOT3D (top), ARCTIC (middle), and EgoDex (bottom), with mesh overlays and trajectory visualizations. Extensive visualization results can be found on the project page.
Figure 4: Qualitative comparison of contact and force estimation. We visualize the contact (left) and force (right) estimation from OpenTouch (first to third row) and PVDB (fourth row).
Test set
Method
F Acc ↑
MPJPE-p ↓
PA-p ↓
MPJPE +OOS ↓
EPE2D-p ↓
Jitter ↓
σshape↓
HOT3D
HaMeR ( Pavlakos et al., 2024 )
0.743
65.53
11.42
66.85
130.33
25.40
1.17
WiLoR ( Potamias et al., 2024 )
0.743
44.89
9.71
46.23
134.89
22.17
1.90
HaWoR ( Zhang et al., 2025 )
0.743
41.53
10.15
42.99
126.10
26.13
7.37
HaPTIC ( Ye et al., 2025b )
0.743
62.42
11.26
63.70
130.40
38.27
3.20
HandFlow ( Xu et al., 2026 )
0.729
39.90
9.50
24.87
127.85
16.99
3.68
ACE-Ego-Hand ( Liu et al., 2026 )
0.928
24.41
8.60
20.30
38.09
6.48
1.22
Table 1: Quantitative comparison of motion reconstruction. RLHND outperforms baselines on nearly every metric on every benchmark consistently. We denote best and second best values.
Contact (vertex-level)
Force (kPa)
OpenTouch
DexYCB
HOT3D
ARCTIC
OpenTouch
PVDB
Method
F1 ↑
AUROC ↑
F1 ↑
AUROC ↑
F1 ↑
AUROC ↑
F1 ↑
AUROC ↑
MAE ↓
RMSE ↓
MAE ↓
RMSE ↓
PressureVision ( Grady et al., 2022 )
–
–
–
–
–
–
–
–
1.930
6.190
0.471
3.726
PressureVision++ ( Grady et al., 2024 )
–
–
–
–
–
–
–
–
1.920
6.200
0.248
3.025
HACO ( Jung & Lee, 2025 )
0.373
0.541
0.543
0.883
0.244
0.749
0.580
0.907
–
–
–
–
HOPE ( Jeon et al., 2026 )
0.663
0.894
0.506
0.868
0.197
0.762
0.591
0.914
1.781
5.388
0.236
2.091
Table 2: Quantitative comparison of contact and force estimation. RLHND outperforms baselines on nearly every benchmark consistently. We denote best and second best values.
Sharpa Wave (22 DoF)
WUJI v2 (20 DoF)
Shadow (24 DoF)
Inspire RH56 (12 DoF)
ALLEX (15 DoF)
Method
Q-err ↓
Jerk ↓
Q-err ↓
Jerk ↓
Q-err ↓
Jerk ↓
Q-err ↓
Jerk ↓
Q-err ↓
Jerk ↓
Oracle (GT joints)
0.0
0.34
0.0
0.33
0.0
0.35
0.0
0.18
0.0
0.36
HaMeR ( Pavlakos et al., 2024 )
16.6
1.45
19.4
1.31
12.6
1.31
8.1
0.90
14.1
1.19
WiLoR ( Potamias et al., 2024 )
14.2
1.48
17.0
1.34
10.2
1.20
6.3
0.90
11.2
1.18
HaWoR ( Zhang et al., 2025 )
15.0
0.61
18.0
0.47
12.2
0.50
8.9
0.39
14.4
0.50
HaPTIC ( Ye et al., 2025b )
14.9
1.39
17.6
1.27
11.9
1.10
7.0
0.87
12.7
1.21
Table 3: Retargeting quality on five dexterous right hands (DexPilot, wrist-relative targets), averaged over the HOT3D, ARCTIC ego and EgoDex test sets; per-dataset results are in Appendix B.4 . Q-err is the mean absolute deviation ( ∘ ) of the joint command from the one the same solver produces on the ground-truth joints, and Jerk the median frame-to-frame command noise ( ∘ /frame 2 ); see Appendix A.4 . The grey Oracle row retargets the ground-truth joints.
Method
Human : Robot episodes
Pick and Place
Bimanual
All ↑
Bottle
Box
Ball
Bird
Avg. ↑
Plastic bag
Assemble tissue
Avg. ↑
DPP ( Kim et al., 2026a )
100 : 100
93.8
100.0
81.3
81.3
89.1
53.1
75.0
64.1
80.8
DPP + RLHND (Ours)
100 : 100
93.8
100.0
87.5
93.8
93.8
59.4
90.6
75.0
87.5
Table 4: Real-robot success rate of Dexterous Point Policy. DPP is trained on human demonstrations labeled by its original tracker, and DPP + RLHND on the same demonstrations labeled by RLHND.
ARCTIC ego
EgoDex
Pose variant
MPJPE-p ↓
PA-p ↓
EPE2D-p ↓
Q-err ↓
MPJPE-p ↓
PA-p ↓
EPE2D-p ↓
Q-err ↓
A0: w/o Cosmos 3
16.62
7.55
13.40
12.48
21.57
10.63
22.45
15.64
A1: w/o constrained pose
14.26
6.52
12.33
10.85
19.91
10.38
20.76
15.31
A2: w/o β -conditioning
14.87
6.72
12.27
10.48
20.06
10.33
20.95
15.29
A3: Full model
13.36
6.42
12.19
10.12
19.90
10.12
20.72
15.22
A4: Full model + GT β
13.15
6.45
12.21
9.69
–
–
–
–
Table 5: Leave-one-out ablations of RLHND. Top (pose): variants A0-A4, evaluated on ARCTIC ego and EgoDex. Bottom (tactile): variants B0–B3, evaluated on contact benchmarks and OpenTouch.
Appendix figures & tables17 assets
Supplementary material from the paper’s appendix.
Appendix
Figure 5: β -cache inference on a test recording. Left hand of a HOT3D Aria test recording decoded in 81 -frame windows. Seven evenly spaced frames are shown per window, with the strip below marking frame-level presence or absence and triangles indicating the displayed frames. Windows whose visibility stays under Tv are held back . The left hand first exceeds Tv in window 4, so its shape becomes βcache and windows 1–3 are decoded again; the right hand is cached from window 1.
Figure 6: Contact label differences between distance threshold and signed distance. Contact labels from the conventional unsigned distance (middle) and our signed distance (right) on two DexYCB grasps. We denote the contacted vertices with the red mark.
DexYCB
ARCTIC
Step rule
λ
Iters
Median ↓
P99 ↓
Max ↓
Median ↓
P99 ↓
Max ↓
Cost ↓
None (projection only)
–
–
1.03
2.34
3.23
1.98
3.96
7.85
0.1×
First-order descent
–
3
1.01
2.32
3.20
1.95
3.91
7.80
1×
Cyclic coordinate descent
–
3
0.42
0.77
1.43
0.91
1.54
4.41
3×
Root-to-tip sequential
10−3
3
0.47
0.80
1.43
0.93
1.46
4.34
8×
Gauss–Newton
0
3
0.34
1.27
9.88
0.72
13.21
24.96
1×
Appendix
Table 6: Constrained-label fitting. Label error (mm), i.e . , the mean per-joint distance between the constrained pose and the original label over the 15 hand joints and the 5 fingertips, on 4057 DexYCB and 4089 ARCTIC frames. Every refinement starts from the same projection and runs the same number of iterations, so the rows differ only in the step rule. The undamped solve attains a competitive median but its worst frames degrade by an order of magnitude, which is the failure mode that matters when the output is a training label; damping bounds it. Cost is wall-clock time relative to ours on the same frames and hardware. We denote best and second best values.
Term of Eq. ( 7 )
Component
Weight
Value
Lrot
geodesic rotation error
λgeo
1
Frobenius rotation error
λmse
1
Lβ
shape
λβ
0.1
L3D
wrist-relative joints
λrel
10
camera-frame joints
λcam
5
camera-frame wrist
λw
2
Appendix
Table 7: Stage-1 loss weights. Each row is one component of the corresponding term of Eq. ( 7 ).
Figure 7: Examples of the DPP training data. Top : human demonstrations for the six tasks, e.g . , ball, bottle, box, bird, plastic bag, and tissue, across approach, grasp, hold, and release phases, viewed from the shared ego camera. Gray and teal denote the original DPP, i.e . , HaWoR , and RLHND hand keypoints, respectively. Marker size and fill indicate the corresponding contact labels. Bottom : teleoperated robot episodes with forward-kinematics keypoints in gray .
Figure 8: Automatic contact labels from RLHND. We denote vertices whose predicted contact probability exceeds 0.5 as teal . The strips below each row show the frame-level contact labels for the right (R) and left (L) hands, obtained by thresholding the mean contact probability over the fingertip vertices at 0.5 . Triangles mark the displayed frames and percentages indicate the fraction of frames labeled as contact.
Figure 9: Real-robot platform. An RB-Y1 bimanual robot with a WUJI dexterous hand on each arm in the tabletop workspace used for the real-robot experiments. The insets on the right show close-ups of the two hands, with the robot’s left and right hands outlined in blue and red , respectively.
Figure 10: Real-robot rollouts of DPP + RLHND. Five snapshots of one rollout per task of the DPP + RLHND policy on the RB-Y1 with WUJI hands, from a third-person view. The four pick-and-place tasks (bottle, box, ball, bird) have the right hand grasp the object and drop it into the white container, and the two bimanual tasks are the ones where the left hand picks up and flips the plastic bag before the right hand passes it on, and the left hand removes the used tissue core before the right hand places the new roll on the holder.
Figure 11: Fine-tuning of DPP and DPP + RLHND. Top : action flow-matching loss over 60 k fine-tuning steps. Bottom : keypoint displacement error of sampled actions at every 10 k-step snapshot, measured as the mean L2 distance in cm between predicted and recorded 16 -step displacements of the 21 keypoints. Results are shown for three human : robot ratios; only the 100 : 100 policies are deployed in Table 4 . Gray denotes DPP with its original human labels and teal denotes DPP + RLHND.
Figure 12: Wrist depth of the human hand labels against the stereo depth. The top two rows show the right-wrist depth over one demonstration each of the plastic-bag and tissue tasks, for the stereo target , raw HaWoR , and raw RLHND , before the per-frame alignment. The bottom row shows the per-episode median absolute deviation from the stereo target over all 200 demonstrations, for both trackers before and after the alignment.
HOT3D
ARCTIC ego
EgoDex
Method
Jitter
Jitter z
Jitter xy
Jitter
Jitter z
Jitter xy
Jitter
Jitter z
Jitter xy
HaMeR
25.40
15.84
17.51
21.73
18.21
9.56
15.48
11.88
8.53
WiLoR
22.17
13.45
15.50
25.92
22.84
9.53
13.05
9.71
7.43
HaWoR
26.13
16.96
17.42
29.34
26.12
10.56
16.94
13.83
8.35
HaPTIC
38.27
27.02
23.97
51.09
47.28
16.40
25.81
21.88
12.06
HandFlow
16.99
8.41
13.37
13.47
7.88
9.49
10.89
5.69
8.32
Appendix
Table 8: Jitter decomposition. Jitter (mm/frame 2 ) is split into its depth component Jitter z and image-plane component Jitter xy under the protocol of Table 1 .
Method
Params (M)
Detector (ms)
Model (ms)
Total (ms/frame)
FPS
MPJPE-p (HOT3D)
HaMeR
699
2.7
11.7
14.4
69.4
65.53
WiLoR
668
2.7
11.9
14.6
68.4
44.89
HaWoR
720
2.7
12.1
14.8
67.7
41.53
HaPTIC
1,357
2.7
23.8
26.5
37.7
62.42
HandFlow
866
–
17.4
17.4
57.4
39.90
ACE-Ego-Hand
3,044
–
1.8
1.8
568.6
24.41
Appendix
Table 9: Inference cost. Inference time measured on one NVIDIA H100 (80 GB) over the same HOT3D Aria test windows, using GPU-synchronized model-forward time per video frame with both hands. Video decoding is excluded and the first forward pass is discarded as warm-up. For crop-based trackers, the shared hand-detector cost is reported separately and included in the total. MPJPE-p is from Table 1 .
Figure 13: Throughput against accuracy. Time per frame with both hands on one H100 (ms, linear, axis reversed so that faster is to the right) against MPJPE-p on HOT3D (log scale, lower is better, so better trackers lie further from the origin); bubble area is proportional to the number of parameters used at inference (labeled in billions). Crop-based per-frame trackers include the shared hand detector. Numbers in Table 9 .
Sharpa Wave (22 DoF)
WUJI v2 (20 DoF)
Shadow (24 DoF)
Inspire RH56 (12 DoF)
ALLEX (15 DoF)
Test set
Method
Q-err ↓
Jerk ↓
Q-err ↓
Jerk ↓
Q-err ↓
Jerk ↓
Q-err ↓
Jerk ↓
Q-err ↓
Jerk ↓
HOT3D
Oracle (GT joints)
0.0
0.27
0.0
0.22
0.0
0.27
0.0
0.13
0.0
0.25
HaMeR ( Pavlakos et al., 2024 )
15.0
1.50
17.8
1.46
13.1
1.30
7.1
0.99
13.3
1.36
WiLoR ( Potamias et al., 2024 )
11.8
1.25
14.6
1.24
9.4
1.03
4.9
1.00
9.6
1.11
HaWoR ( Zhang et al., 2025 )
13.6
0.37
15.8
0.32
10.7
0.36
7.2
0.25
11.2
0.32
HaPTIC ( Ye et al., 2025b )
15.3
1.60
18.2
1.52
12.7
1.09
6.8
1.10
13.7
1.66
Appendix
Table 10: Per-dataset retargeting quality on five dexterous right hands (DexPilot, wrist-relative targets); Table 3 in the main text reports the mean of the three blocks. Q-err is the mean absolute deviation ( ∘ ) of the joint command from the one the same solver produces on the ground-truth joints; Jerk is the median over frames of the joint-averaged second temporal difference of the command ( ∘ /frame 2 ), i.e. the frame-to-frame noise on typical frames (Appendix A.4 ). The grey Oracle row retargets the ground-truth joints (Q-err =0 by construction). We denote best and second best values with shade and bold.
Figure 14: Extensive qualitative comparison of 2D and 3D hand motion reconstruction.
Figure 15: Extensive qualitative comparison of per-vertex contact and force estimation.
OpenTouch
DexYCB
HOT3D
OpenTouch force (kPa)
Variant
F1 ↑
AUROC ↑
F1 ↑
AUROC ↑
F1 ↑
AUROC ↑
MAE ↓
RMSE ↓
w/ one-way attention
0.635
0.979
0.565
0.910
0.593
0.955
0.573
2.574
Ours (no pose conditioning)
0.696
0.980
0.572
0.915
0.589
0.959
0.489
2.508
Appendix
Table 11: One-way attention ablation. Comparison between the proposed tactile expert and a variant in which the bone tokens additionally attend to the pose stream. Best values are shaded.
We introduce a novel 3D hand pose estimator that can accurately recover the shape and pose of people's hands in a room from afar, typically from fixed cameras at room corners, in extremely low-resolution and frequently occluded views. Our key idea is to fully leverage hand-body coordination, its temporal progression, and multiview observations. We achieve this with a novel Transformer-based model, in which hand and body configurations are modeled through correlations between their visual features expressed as per-view tokens, and their temporal coordination is exploited in an autoregressive manner. We introduce a novel dataset, which we refer to as REACH, Room-Environment dataset Annotated with Chest cameras for Hand pose estimation, to train and test our method. REACH is a first-of-its-kind large-scale hand pose dataset that captures accurate hand movements of 50 participants across a wide variety of daily activities. In order to avoid interfering with natural movements while annotating the hands with accurate shape and pose, we leverage concealed chest cameras. Through extensive experiments, including comparative studies with existing methods, we show that our model, REACH-Net, achieves highly accurate 3D hand pose estimation from afar. These results broaden the horizon of 3D hand pose estimation, especially towards "in-the-wild" continuous human behavior analysis.
Shu Nakamura, Ryo Kawahara, Genki Kinoshita +4
Graduate School of Informatics, Kyoto University · RIKEN · Kyoto Institute of Technology
Reconstructing the motion of objects from videos is a key component for embodied AI and robot manipulation. While diverse approaches to object pose tracking have been studied, they rely heavily on strong external priors, such as depth data or 3D templates, and remain highly vulnerable to severe occlusions by hand grasps despite the use of explicit masks. In this work, we present ComPose, a 6DoF object tracking framework designed for hand-aware object pose estimation from RGB video. Rather than treating the hand purely as an occluder, our method harmonizes hand motions as a \textit{complementary cue} for object tracking. In detail, we recover a variety of object motions over time by combining object and hand cues from foundation models within a unified tracking pipeline. Here, ComPose adaptively selects informative hand joints, combines object- and hand-derived cues for motion estimation, and refines the resulting object motion using visible geometric evidence and a learned correction. We further enforce the temporal consistency over both rotation and translation, yielding stable 3D object trajectories over time without any external smoothing. Extensive experiments show that our method is accurate, efficient, and robust under severe hand occlusion and geometric ambiguity. In addition, the resulting trajectories can also effectively transfer to downstream robot manipulation by enabling robots to reconstruct human actions from online videos.
Human-hand demonstrations provide a direct and scalable source of physical interaction data for robot learning. While manual retargeting is indispensable for establishing kinematic action correspondence across different morphologies, robust transfer requires going beyond geometry to address the underlying alignment of physical dynamics between human and robot manipulation. To address this, we introduce LaST-HD, a novel human-to-robot action learning paradigm that extends reasoning-before-acting VLA by aligning human-hand and robot demonstrations in a shared latent reasoning space. Rather than mimicking human kinematics, LaST-HD trains an auxiliary action-conditioned world model on unpaired human-hand and robot trajectories to synthesize unified latent targets. After aligning cross-embodiment representations in this shared forward-dynamics space, these targets supervise LaST-HD's latent reasoning process, enabling it to internalize shared physical dynamics and drive efficient human-hand action learning. Moreover, we develop Out-of-Lab (OOL) Glove, a low-cost motion-capture glove tailored to LaST-HD for human-hand data collection. The captured human data provide precise keypoints and serve as universal action supervision across grippers and dexterous hands. Armed with the aligned latent space and high-fidelity human-hand data, we develop a progressive mixed-to-human training recipe comprising mixed human-robot co-training and human-hand online correction post-training. Through mixed co-training, LaST-HD improves generalization to novel objects, scenes, and positions using only human-hand demonstrations. With online correction, LaST-HD further adapts to novel environments and achieves over 90% accuracy using only 20 minutes of OOL glove data.
Jiaming Liu, Yinxi Wang, Chenyang Gu +15
State Key Laboratory of Multimedia Information Processing, School of Computer Science, Peking University · The Chinese University of Hong Kong · Aether Tech +1