RLHND: Video Foundation Models as Physically Grounded Hand Trackers for Robot Learning
Organizations: RLWRLD · KAIST · Seoul National University
Abstract
Recently, approaches that leverage human video datasets for robot policy training have become increasingly prevalent. However, most existing hand trackers regress pose from cropped frames with limited priors on hand motion and object interaction, resulting in inaccurate and physically inconsistent estimates. Moreover, the lack of physical cues, e.g., contact and force, limits the use of human videos for robot policy training. To this end, we propose RLHND, a video foundation model-based hand tracking model that jointly estimates hand pose and realistic tactile information from monocular egocentric videos. RLHND turns the pre-trained Cosmos 3 video diffusion backbone into a deterministic clip-level feature extractor via clean-latent conditioning, carrying its learned priors on hand motion and hand-object interaction into tracking. For pose estimation, RLHND (i) predicts hand poses with anatomically plausible joint angles and (ii) enables optional conditioning on the shape parameter to maintain consistent hand shape within the same video and even across videos recorded by the same actor. For tactile estimation, a separate tactile expert stream, trained with the pose stream frozen, predicts dense contact and force over the hand surface. We further adopt LBS-based feature spreading to enable vertex-wise feature extraction without costly per-vertex attention. RLHND achieves state-of-the-art performance across various benchmark datasets for pose estimation, while also achieving state-of-the-art performance in contact and force estimation. Moreover, we demonstrate the utility of RLHND for robot learning through retargeting results and real-world robot experiments. The code will be publicly available at https://seungjun-moon.github.io/rlhnd/.
Figures & tables
| Test set | Method | F Acc | MPJPE-p | PA-p | MPJPE +OOS | EPE2D-p | Jitter | |
|---|---|---|---|---|---|---|---|---|
| HOT3D | HaMeR ( Pavlakos et al., 2024 ) | 0.743 | 65.53 | 11.42 | 66.85 | 130.33 | 25.40 | 1.17 |
| WiLoR ( Potamias et al., 2024 ) | 0.743 | 44.89 | 9.71 | 46.23 | 134.89 | 22.17 | 1.90 | |
| HaWoR ( Zhang et al., 2025 ) | 0.743 | 41.53 | 10.15 | 42.99 | 126.10 | 26.13 | 7.37 | |
| HaPTIC ( Ye et al., 2025b ) | 0.743 | 62.42 | 11.26 | 63.70 | 130.40 | 38.27 | 3.20 | |
| HandFlow ( Xu et al., 2026 ) | 0.729 | 39.90 | 9.50 | 24.87 | 127.85 | 16.99 | 3.68 | |
| ACE-Ego-Hand ( Liu et al., 2026 ) | 0.928 | 24.41 | 8.60 | 20.30 | 38.09 | 6.48 | 1.22 |
| Contact (vertex-level) | Force (kPa) | |||||||||||
|---|---|---|---|---|---|---|---|---|---|---|---|---|
| OpenTouch | DexYCB | HOT3D | ARCTIC | OpenTouch | PVDB | |||||||
| Method | F1 | AUROC | F1 | AUROC | F1 | AUROC | F1 | AUROC | MAE | RMSE | MAE | RMSE |
| PressureVision ( Grady et al., 2022 ) | – | – | – | – | – | – | – | – | 1.930 | 6.190 | 0.471 | 3.726 |
| PressureVision++ ( Grady et al., 2024 ) | – | – | – | – | – | – | – | – | 1.920 | 6.200 | 0.248 | 3.025 |
| HACO ( Jung & Lee, 2025 ) | 0.373 | 0.541 | 0.543 | 0.883 | 0.244 | 0.749 | 0.580 | 0.907 | – | – | – | – |
| HOPE ( Jeon et al., 2026 ) | 0.663 | 0.894 | 0.506 | 0.868 | 0.197 | 0.762 | 0.591 | 0.914 | 1.781 | 5.388 | 0.236 | 2.091 |
| Sharpa Wave (22 DoF) | WUJI v2 (20 DoF) | Shadow (24 DoF) | Inspire RH56 (12 DoF) | ALLEX (15 DoF) | ||||||
| Method | Q-err | Jerk | Q-err | Jerk | Q-err | Jerk | Q-err | Jerk | Q-err | Jerk |
| Oracle (GT joints) | 0.0 | 0.34 | 0.0 | 0.33 | 0.0 | 0.35 | 0.0 | 0.18 | 0.0 | 0.36 |
| HaMeR ( Pavlakos et al., 2024 ) | 16.6 | 1.45 | 19.4 | 1.31 | 12.6 | 1.31 | 8.1 | 0.90 | 14.1 | 1.19 |
| WiLoR ( Potamias et al., 2024 ) | 14.2 | 1.48 | 17.0 | 1.34 | 10.2 | 1.20 | 6.3 | 0.90 | 11.2 | 1.18 |
| HaWoR ( Zhang et al., 2025 ) | 15.0 | 0.61 | 18.0 | 0.47 | 12.2 | 0.50 | 8.9 | 0.39 | 14.4 | 0.50 |
| HaPTIC ( Ye et al., 2025b ) | 14.9 | 1.39 | 17.6 | 1.27 | 11.9 | 1.10 | 7.0 | 0.87 | 12.7 | 1.21 |
| Method | Human : Robot episodes | Pick and Place | Bimanual | All | ||||||
|---|---|---|---|---|---|---|---|---|---|---|
| Bottle | Box | Ball | Bird | Avg. | Plastic bag | Assemble tissue | Avg. | |||
| DPP ( Kim et al., 2026a ) | 100 : 100 | 93.8 | 100.0 | 81.3 | 81.3 | 89.1 | 53.1 | 75.0 | 64.1 | 80.8 |
| DPP + RLHND (Ours) | 100 : 100 | 93.8 | 100.0 | 87.5 | 93.8 | 93.8 | 59.4 | 90.6 | 75.0 | 87.5 |
| ARCTIC ego | EgoDex | |||||||
| Pose variant | MPJPE-p | PA-p | EPE2D-p | Q-err | MPJPE-p | PA-p | EPE2D-p | Q-err |
| A0: w/o Cosmos 3 | 16.62 | 7.55 | 13.40 | 12.48 | 21.57 | 10.63 | 22.45 | 15.64 |
| A1: w/o constrained pose | 14.26 | 6.52 | 12.33 | 10.85 | 19.91 | 10.38 | 20.76 | 15.31 |
| A2: w/o -conditioning | 14.87 | 6.72 | 12.27 | 10.48 | 20.06 | 10.33 | 20.95 | 15.29 |
| A3: Full model | 13.36 | 6.42 | 12.19 | 10.12 | 19.90 | 10.12 | 20.72 | 15.22 |
| A4: Full model + GT | 13.15 | 6.45 | 12.21 | 9.69 | – | – | – | – |
Appendix figures & tables17 assets
Supplementary material from the paper’s appendix.
Appendix
| DexYCB | ARCTIC | ||||||||
| Step rule | Iters | Median | P99 | Max | Median | P99 | Max | Cost | |
| None (projection only) | – | – | 1.03 | 2.34 | 3.23 | 1.98 | 3.96 | 7.85 | |
| First-order descent | – | 3 | 1.01 | 2.32 | 3.20 | 1.95 | 3.91 | 7.80 | |
| Cyclic coordinate descent | – | 3 | 0.42 | 0.77 | 1.43 | 0.91 | 1.54 | 4.41 | |
| Root-to-tip sequential | 3 | 0.47 | 0.80 | 1.43 | 0.93 | 1.46 | 4.34 | ||
| Gauss–Newton | 3 | 0.34 | 1.27 | 9.88 | 0.72 | 13.21 | 24.96 | ||
| Term of Eq. ( 7 ) | Component | Weight | Value |
|---|---|---|---|
| geodesic rotation error | 1 | ||
| Frobenius rotation error | 1 | ||
| shape | 0.1 | ||
| wrist-relative joints | 10 | ||
| camera-frame joints | 5 | ||
| camera-frame wrist | 2 |
| HOT3D | ARCTIC ego | EgoDex | |||||||
| Method | Jitter | Jitter z | Jitter xy | Jitter | Jitter z | Jitter xy | Jitter | Jitter z | Jitter xy |
| HaMeR | 25.40 | 15.84 | 17.51 | 21.73 | 18.21 | 9.56 | 15.48 | 11.88 | 8.53 |
| WiLoR | 22.17 | 13.45 | 15.50 | 25.92 | 22.84 | 9.53 | 13.05 | 9.71 | 7.43 |
| HaWoR | 26.13 | 16.96 | 17.42 | 29.34 | 26.12 | 10.56 | 16.94 | 13.83 | 8.35 |
| HaPTIC | 38.27 | 27.02 | 23.97 | 51.09 | 47.28 | 16.40 | 25.81 | 21.88 | 12.06 |
| HandFlow | 16.99 | 8.41 | 13.37 | 13.47 | 7.88 | 9.49 | 10.89 | 5.69 | 8.32 |
| Method | Params (M) | Detector (ms) | Model (ms) | Total (ms/frame) | FPS | MPJPE-p (HOT3D) |
|---|---|---|---|---|---|---|
| HaMeR | 699 | 2.7 | 11.7 | 14.4 | 69.4 | 65.53 |
| WiLoR | 668 | 2.7 | 11.9 | 14.6 | 68.4 | 44.89 |
| HaWoR | 720 | 2.7 | 12.1 | 14.8 | 67.7 | 41.53 |
| HaPTIC | 1,357 | 2.7 | 23.8 | 26.5 | 37.7 | 62.42 |
| HandFlow | 866 | – | 17.4 | 17.4 | 57.4 | 39.90 |
| ACE-Ego-Hand | 3,044 | – | 1.8 | 1.8 | 568.6 | 24.41 |
| Sharpa Wave (22 DoF) | WUJI v2 (20 DoF) | Shadow (24 DoF) | Inspire RH56 (12 DoF) | ALLEX (15 DoF) | |||||||
| Test set | Method | Q-err | Jerk | Q-err | Jerk | Q-err | Jerk | Q-err | Jerk | Q-err | Jerk |
| HOT3D | Oracle (GT joints) | 0.0 | 0.27 | 0.0 | 0.22 | 0.0 | 0.27 | 0.0 | 0.13 | 0.0 | 0.25 |
| HaMeR ( Pavlakos et al., 2024 ) | 15.0 | 1.50 | 17.8 | 1.46 | 13.1 | 1.30 | 7.1 | 0.99 | 13.3 | 1.36 | |
| WiLoR ( Potamias et al., 2024 ) | 11.8 | 1.25 | 14.6 | 1.24 | 9.4 | 1.03 | 4.9 | 1.00 | 9.6 | 1.11 | |
| HaWoR ( Zhang et al., 2025 ) | 13.6 | 0.37 | 15.8 | 0.32 | 10.7 | 0.36 | 7.2 | 0.25 | 11.2 | 0.32 | |
| HaPTIC ( Ye et al., 2025b ) | 15.3 | 1.60 | 18.2 | 1.52 | 12.7 | 1.09 | 6.8 | 1.10 | 13.7 | 1.66 | |
| OpenTouch | DexYCB | HOT3D | OpenTouch force (kPa) | |||||
|---|---|---|---|---|---|---|---|---|
| Variant | F1 | AUROC | F1 | AUROC | F1 | AUROC | MAE | RMSE |
| w/ one-way attention | 0.635 | 0.979 | 0.565 | 0.910 | 0.593 | 0.955 | 0.573 | 2.574 |
| Ours (no pose conditioning) | 0.696 | 0.980 | 0.572 | 0.915 | 0.589 | 0.959 | 0.489 | 2.508 |