cs.CVOct 7, 2026

RLHND: Video Foundation Models as Physically Grounded Hand Trackers for Robot Learning

Authors: Seungjun Moon, Subin Jeon, Sangwoo Kim, Hanbyul Joo, Jinwoo Shin

Organizations: RLWRLD · KAIST · Seoul National University

Abstract

Recently, approaches that leverage human video datasets for robot policy training have become increasingly prevalent. However, most existing hand trackers regress pose from cropped frames with limited priors on hand motion and object interaction, resulting in inaccurate and physically inconsistent estimates. Moreover, the lack of physical cues, e.g., contact and force, limits the use of human videos for robot policy training. To this end, we propose RLHND, a video foundation model-based hand tracking model that jointly estimates hand pose and realistic tactile information from monocular egocentric videos. RLHND turns the pre-trained Cosmos 3 video diffusion backbone into a deterministic clip-level feature extractor via clean-latent conditioning, carrying its learned priors on hand motion and hand-object interaction into tracking. For pose estimation, RLHND (i) predicts hand poses with anatomically plausible joint angles and (ii) enables optional conditioning on the shape parameter to maintain consistent hand shape within the same video and even across videos recorded by the same actor. For tactile estimation, a separate tactile expert stream, trained with the pose stream frozen, predicts dense contact and force over the hand surface. We further adopt LBS-based feature spreading to enable vertex-wise feature extraction without costly per-vertex attention. RLHND achieves state-of-the-art performance across various benchmark datasets for pose estimation, while also achieving state-of-the-art performance in contact and force estimation. Moreover, we demonstrate the utility of RLHND for robot learning through retargeting results and real-world robot experiments. The code will be publicly available at https://seungjun-moon.github.io/rlhnd/.

Figures & tables

Appendix figures & tables17 assets

Supplementary material from the paper’s appendix.

Appendix

Explore similar work

CardsList
  1. REACH: Hand Pose Estimation from Room Corners

    May 21, 2026Shu Nakamura, Ryo Kawahara, Genki Kinoshita +43D Hand

  2. ComPose: When to Trust Hands for Object Pose Tracking

    May 22, 2026Jisu Shin, Junoh Lee, JunGyu Lee +53D HandCategory-Level Object Pose Estimation

  3. LaST-HD: Learning Latent Physical Reasoning from Scalable Human Data for Robot Manipulation

    Jun 22, 2026Jiaming Liu, Yinxi Wang, Chenyang Gu +15Human-To-Robot TransferRobotic Manipulation