Learning dexterous manipulation benefits from human demonstration datasets that capture diverse and natural hand-object interactions. In particular, contact points and forces provide supervision on where and how strongly to interact, which cannot be fully captured by motion trajectories alone. However, methods for jointly capturing hand and object motion, contact points, and forces remain limited. Moreover, severe occlusion from hand-object interaction challenges accurate tracking of both hands and objects. To address these limitations, we present a Visual-Inertial-CONtact based hand-object tracking (VICON) framework. It holistically captures both hand and object motion along with contact information during manipulation, even under severe occlusion. First, we adopt a visual-inertial glove and an RGB-D camera for accurate hand tracking, and redesign the glove to incorporate contact sensing. Specifically, force-sensitive resistors (FSRs) are placed on the glove based on human grasp frequency to synchronously record contact states and calibrated normal forces. Second, without requiring pre-existing CAD models, we estimate object poses using RGB-D images and a mesh reconstructed from a monocular video. We propose factor-graph-based object trajectory estimation that fuses object-pose estimates weighted by visibility under hand-object occlusion, FSR measurements, and a hand-motion prior. Across 40 motion-capture sessions with five objects, VICON achieves a 2.5% failed-frame rate compared with 50.9-64.6% for the baselines, with median errors of 3.9 mm and 3.0 degrees under occlusion. Using VICON, we construct a dataset containing synchronized hand-object motion, contact points, and normal forces, and will publicly release an expanded dataset covering 10 object categories at https://github.com/VICON-dataset/dataset.
Figures & tables
Fig. 1: VICON records hand motion, object pose, contact points, and forces using a visual-inertial glove [ 1 ] with integrated FSRs and a single RGB-D camera. The glove tracks the hand skeleton, and our offline refinement maintains robust and accurate object 6-DoF tracking even under severe occlusion. Inset: object-frame contact points, with arrows colored by calibrated normal force magnitude.
Dataset
Hand
Object
Contact
Force/
Setup
pose
pose
format
tactile
EgoDex [ 2 ]
✓
–
–
–
Vision Pro
ContactPose [ 6 ]
✓
✓
Object surface map
–
3 RGB-D, thermal, mocap
ARCTIC [ 7 ]
✓
✓
Object surface map
–
8+1 RGB, mocap
OpenTouch [ 9 ]
✓
–
Hand map
Pressure
Aria glasses, tracking and tactile gloves
TactiDex [ 10 ]
✓
✓
Hand map
Pressure
Mocap, tracking and tactile gloves
TABLE I: Comparison of our dataset with existing interaction datasets
Fig. 2: Multi-layer VICON-glove. (a) Placement of the 14 IMUs in the inner glove. (b) Outer glove with the visual markers. (c) Hand contact-frequency map [ 17 ] , averaged over the mechanic and housemaid tasks (values in %), used to choose the FSR sites. (d) Palmar view of the eight FSRs.
Fig. 3: VICON on a recorded session. The VICON-glove provides hand tracking (poses H1 – H5 , increments ΔH ) and contact (FSR); the RGB-D camera feeds FoundationPose, whose candidates enter as fvis (filled: accepted; hollow: rejected under occlusion). Between object poses Xk , contact selects fstat (visible, after release; red crosses: rejected hand increments), fpivot (2-point grasp), or fhand (3-point grasp). Re-measurement feeds the refined poses back to FoundationPose.
vd
0.10
0.15
0.30
0.50
0.70
0.90
1.0
σp (mm)
6
5
3
3
2
1.2
0.8
σr (deg)
180
180
8
3
2
1
0.8
TABLE II: Nominal visual-measurement uncertainty as a function of visibility
Fig. 4: Experimental setup. (a) Motion-capture room; the inset shows one of the motion-capture cameras. (b) The five objects: block, mouse, mug, mustard bottle, and phone.
Fig. 5: Object-pose errors over the 40 motion-capture sessions (top: translation; bottom: symmetry-aware rotation). (a) Median error of each tracker binned by the visibility vdGT ; shading marks the occluded range. (b) PCF on the 24,285 occluded frames as a function of the error threshold; dotted lines mark the 20 mm and 15 ∘ thresholds used for ST20 and SR15 , and the numbers give the AUC. Medians are used in (a) because diverging baseline errors would dominate a mean.
Fig. 6: Qualitative results with increasing hand occlusion from top to bottom. Columns: object, 3DGS mesh, and the input image overlaid with the ground-truth pose (green) and the pose of FoundationPose, ICG+, and VICON. Last column: rotated view of the tracked hand, visualized as a MANO [ 27 ] mesh, and the VICON object pose with the contact points and force arrows of the recognized FSRs, colored by the calibrated normal force.
Robotics Research Center (RRC), International Institute of Information Technology (IIIT), Hyderabad, India · Department of Computer Science, The University of Manchester