Functional Hand Type Prior for 3D Hand Pose Estimation and Action Recognition from Egocentric View Monocular Videos
Authors: Wonseok Roh, Seung Hyun Lee, Won Jeong Ryoo, Jakyung Lee, Gyeongrok Oh, Sooyeon Hwang, Hyung-gun Chi, Sangpil Kim
Organizations: Department of Artificial Intelligence, Korea University, Seoul, Republic of Korea · Electrical and Computer Engineering, Purdue University, West Lafayette, Indiana, USA
Current methods for egocentric view action recognition often face challenges in perceiving dynamic hand movements relying solely on geometrical or physical information. In this work, we effectively address this problem by gaining insights into the correlation between functional hand configurations and objects, which improves the detailed interpretation of real-world scenarios. To this end, we introduce a practical taxonomy of hand types based on the functioning perspective and utilize it for per-frame hand type labeling on existing datasets. We also propose a novel hand action recognition framework considering semantic details of the hand type as prior. This approach boosts the network's understanding of the continuous hand interaction throughout the action sequence. Our whole pipeline consists of three main modules: (1) Feature Extraction, (2) Egocentric Knowledge Module, which estimates 3D hand pose, object category, and hand type leveraging short-term cues, and (2) Egocentric Action Module, which aggregates per-frame knowledge, including text embeddings of hand type, over a longer time. In our extensive experiments with large-scale benchmarks, FPHA and H2O, our model outperforms current state-of-the-art methods, demonstrating its superior performance.
Figures & tables
Figure 1: Examples of actions and corresponding hand types from the FPHA dataset [ 15 ] . While each action shares the goal of opening objects, the specific interactions between hands and each object are different. Further, both activities involve holding and opening the lid, but the semantic functions of these motions differ significantly. To highlight the continuously repeated shape during the operation, we represent it using “Dynamic” .
Figure 2: Taxonomy of hand types based on functionality. We categorize the mainly used hand types focusing on the role of the hand. Beyond static hand grasp types, we further define dynamic hand types to Control to explain more real-life hand behaviors.
Figure 3: Overview of our proposed model. Our framework consists of three main modules: (1) Feature Extraction, (2) Egocentric Knowledge Module, which estimates 3D hand pose, object category and hand type leveraging short-term temporal cues, and (3) Egocentric Action Module, which aggregates per-frame pose, object and hand type information over a longer time span.
Joule-color [ 17 ]
Two Stream [ 10 ]
H+O [ 44 ]
Collaborative [ 52 ]
HTT [ 49 ]
Trear [ 25 ]
Ours
Accuracy ( ↑ )
66.78
75.30
82.43
85.22
94.09
94.96
95.13
Table 1: Comparison of our novel hand action recognition framework and the state-of-the-art models on the FPHA [ 15 ] and H2O [ 24 ] dataset. We report the classification accuracy of methods based on RGB videos. Note that the H2O dataset provides additional testing split videos, unlike the FPHA dataset, which provides only training and validation split.
Model
AUC-RA(0-50) ( ↑ )
MEPE-RA ( ↓ )
HTT [ 49 ]
0.763
12.13
Ours
0.769
11.79
Table 2: 3D pose estimation performance in Root-Aligned space on the FPHA [ 15 ] and H2O [ 24 ] . We report AUC-RA for 3D PCK-RA at error thresholds ranging from 0 to 50 mm and the MEPE-RA in the unit of mm .
Hand Type
EAM Input
Text Embedding
Accuracy ( ↑ )
FPHA [ 15 ]
H2O [ 24 ]
✓
-
-
93.74
85.95
✓
✓
-
94.26
87.60
✓
✓
✓
95.13
89.67
Table 3: Ablative study of input features for Egocentric Action Module (EAM) on Hand Action Recognition Accuracy (%). We investigate the usage of the hand type feature in (a). Also, we analyze the effectiveness of each cue on the action recognition task in (b).
Figure 4: Qualitative result of our experiments. In (a), the green and blue line represents ground truth and estimated 3D hand pose, respectively. (b) shows the 3D PCK of hand pose estimation results on H2O [ 24 ] in Root-Aligned space. The blue line indicates the performance of our model, whereas the red line represents the HTT [ 49 ] .
Appendix figures & tables5 assets
Supplementary material from the paper’s appendix.
Appendix
Figure 1: Qualitative comparison of 3D hand pose estimation results between our model (blue lines) and HTT (magenta lines) on the (a) FPHA [ 15 ] and (b) H2O [ 24 ] dataset. Both ground truth (green lines) and estimated hand poses are visualized in 3D space and projected to the corresponding 2D image frames. Note that the H2O provides labels for both hands, while the FPHA only includes information for the right hand.
Figure 2: Examples of action scenes in which the variation of hand pose is similar, but the action labels are different according to the temporal context of hand function and object types. (a) “Index Finger” , (b) “Poking” , and (c) “Extended Index Curl” are all hand types that use index fingers, but their functions vary depending on the action and object they interact with. “Index Finger” represents holding an object while extending the index finger, “Poking” represents poking into an object without deforming it, and “Extended Index Curl” represents utilizing the index finger to deform an object by applying pressure on its tip.
Figure 3: Examples of continuous video frames with hand type annotation. The hand type of each frame changes over time in the video sequence. In (a) and (b), we show frames of the “Place Cappuccino” action (H2O [ 24 ] ) in which the hand interacts with the box and the “High Five” action (FPHA [ 15 ] ) in which the hand interacts with the other’s hand, respectively.
Figure 5: Hand Type Distributions of the (a) FPHA [ 15 ] Dataset (Orange) and (b) H2O [ 24 ] Dataset (Blue). The histograms (left) show the number of frames per hand type; the x-axis represents the hand type, and the y-axis indicates the number of frames. The scatter plots (right) present the distribution of hand types across each action label; the x-axis represents the action of each dataset, and the y-axis indicates the hand type. The size of the circles illustrates how often each hand type appears within the hand action video sequences.
Figure 6: Hand Type Distributions for the (a) Left hand and (b) Right hand of the H2O [ 24 ] Dataset. The histograms (left) show the number of frames per hand type; the x-axis represents the hand type, and the y-axis indicates the number of frames. The scatter plots (right) present the distribution of hand types across each action label; the x-axis represents the action of each dataset, and the y-axis indicates the hand type. The size of the circles illustrates how often each hand type appears within the hand action video sequences.
Forecasting future 3D hand pose sequences from egocentric video is essential for understanding human intention and enabling embodied applications such as AR/VR assistance and human-robot interaction. However, this task remains a highly challenging problem because egocentric hand motion is driven by complex human intent, exhibits highly dexterous articulations, and is observed under drastic viewpoint shifts induced by ego-motion. In this work, we introduce EggHand, a foundation-model-based framework for egocentric hand pose forecasting that unifies multimodal semantic reasoning with dynamic motion modeling. Our approach couples an action decoder from a Vision-Language-Action (VLA) model, which captures the structured temporal dynamics of hand motion, with an egocentric video-text encoder that provides viewpoint-aware contextual information learned from large-scale first-person video. Together, these components overcome the brittleness of generic visual encoders under ego-motion and enable joint reasoning over motion, context, and high-level intent-without relying on body pose or external tracking. Experiments on the EgoExo4D dataset show that EggHand sets a new state of the art in forecasting accuracy, remains robust under severe ego-motion, and further enables controllable prediction via language-based task prompts. Project page: https://jyoun9.github.io/EggHand
Reconstructing the absolute 3D pose and shape of the hands from the user's viewpoint using a single head-mounted camera is crucial for practical egocentric interaction in AR/VR, telepresence, and hand-centric manipulation tasks, where sensing must remain compact and unobtrusive. While monocular RGB methods have made progress, they remain constrained by depth-scale ambiguity and struggle to generalize across the diverse optical configurations of head-mounted devices. As a result, models typically require extensive training on device-specific datasets, which are costly and laborious to acquire. This paper addresses these challenges by introducing EgoForce, a monocular 3D hand reconstruction framework that recovers robust, absolute 3D hand pose and its position from the user's (camera-space) viewpoint. EgoForce operates across fisheye, perspective, and distorted wide-FOV camera models using a single unified network. Our approach combines a differentiable forearm representation that stabilizes hand pose, a unified arm-hand transformer that predicts both hand and forearm geometry from a single egocentric view, mitigating depth-scale ambiguity, and a ray space closed-form solver that enables absolute 3D pose recovery across diverse head-mounted camera models. Experiments on three egocentric benchmarks show that EgoForce achieves state-of-the-art 3D accuracy, reducing camera-space MPJPE by up to 28% on the HOT3D dataset compared to prior methods and maintaining consistent performance across camera configurations. For more details, visit the project page at https://dfki-av.github.io/EgoForce.
Christen Millerdurai, Shaoxiang Wang, Yaxu Xie +3
Deutsches Forschungszentrum für Künstliche Intelligenz (DFKI), Kaiserslautern, Germany · Max Planck Institute for Informatics (MPII), Saarbrücken, Germany
Estimating accurate 3D hand-object pose from in-the-wild egocentric RGB remains challenging due to severe occlusions and ambiguous contact. Existing learning-based methods often struggle to generalise to in-the-wild scenes and are limited by the scarcity of supervision. We address these issues with two contributions. First, we introduce EPIC-Contact, an in-the-wild egocentric dataset of 2.3K clips (62.3K frames) with dense, bijective 3D hand-object contact correspondences and posed meshes. Second, we propose HOPformer, an end-to-end transformer that jointly predicts bi-manual hand and object pose in a single forward pass. A cross-attention decoder conditions object features on hand priors, producing robust pose estimation. We test HOPformer on the in-lab 3D dataset, ARCTIC, as well as our newly introduced EPIC-Contact dataset. HOPformer reaches 82.4% success rate on ARCTIC (+6.2 pts over current SOTA). On EPIC-Contact, it nearly doubles the success rate while reducing contact deviation by 75%. EPIC-Contact, HOPformer code and checkpoints are released: https://sid2697.github.io/epic-contact.
Siddhant Bansal, Zhifan Zhu, Shashank Tripathi +3
University of Bristol, Bristol, United Kingdom · Max Planck Institute for Intelligent Systems, Tübingen, Germany