Learning dexterous manipulation benefits from human demonstration datasets that capture diverse and natural hand-object interactions. In particular, contact points and forces provide supervision on where and how strongly to interact, which cannot be fully captured by motion trajectories alone. However, methods for jointly capturing hand and object motion, contact points, and forces remain limited. Moreover, severe occlusion from hand-object interaction challenges accurate tracking of both hands and objects. To address these limitations, we present a Visual-Inertial-CONtact based hand-object tracking (VICON) framework. It holistically captures both hand and object motion along with contact information during manipulation, even under severe occlusion. First, we adopt a visual-inertial glove and an RGB-D camera for accurate hand tracking, and redesign the glove to incorporate contact sensing. Specifically, force-sensitive resistors (FSRs) are placed on the glove based on human grasp frequency to synchronously record contact states and calibrated normal forces. Second, without requiring pre-existing CAD models, we estimate object poses using RGB-D images and a mesh reconstructed from a monocular video. We propose factor-graph-based object trajectory estimation that fuses object-pose estimates weighted by visibility under hand-object occlusion, FSR measurements, and a hand-motion prior. Across 40 motion-capture sessions with five objects, VICON achieves a 2.5% failed-frame rate compared with 50.9-64.6% for the baselines, with median errors of 3.9 mm and 3.0 degrees under occlusion. Using VICON, we construct a dataset containing synchronized hand-object motion, contact points, and normal forces, and will publicly release an expanded dataset covering 10 object categories at https://github.com/VICON-dataset/dataset.
Figures & tables
Fig. 1: VICON records hand motion, object pose, contact points, and forces using a visual-inertial glove [ 1 ] with integrated FSRs and a single RGB-D camera. The glove tracks the hand skeleton, and our offline refinement maintains robust and accurate object 6-DoF tracking even under severe occlusion. Inset: object-frame contact points, with arrows colored by calibrated normal force magnitude.
Dataset
Hand
Object
Contact
Force/
Setup
pose
pose
format
tactile
EgoDex [ 2 ]
✓
–
–
–
Vision Pro
ContactPose [ 6 ]
✓
✓
Object surface map
–
3 RGB-D, thermal, mocap
ARCTIC [ 7 ]
✓
✓
Object surface map
–
8+1 RGB, mocap
OpenTouch [ 9 ]
✓
–
Hand map
Pressure
Aria glasses, tracking and tactile gloves
TactiDex [ 10 ]
✓
✓
Hand map
Pressure
Mocap, tracking and tactile gloves
TABLE I: Comparison of our dataset with existing interaction datasets
Fig. 2: Multi-layer VICON-glove. (a) Placement of the 14 IMUs in the inner glove. (b) Outer glove with the visual markers. (c) Hand contact-frequency map [ 17 ] , averaged over the mechanic and housemaid tasks (values in %), used to choose the FSR sites. (d) Palmar view of the eight FSRs.
Fig. 3: VICON on a recorded session. The VICON-glove provides hand tracking (poses H1 – H5 , increments ΔH ) and contact (FSR); the RGB-D camera feeds FoundationPose, whose candidates enter as fvis (filled: accepted; hollow: rejected under occlusion). Between object poses Xk , contact selects fstat (visible, after release; red crosses: rejected hand increments), fpivot (2-point grasp), or fhand (3-point grasp). Re-measurement feeds the refined poses back to FoundationPose.
vd
0.10
0.15
0.30
0.50
0.70
0.90
1.0
σp (mm)
6
5
3
3
2
1.2
0.8
σr (deg)
180
180
8
3
2
1
0.8
TABLE II: Nominal visual-measurement uncertainty as a function of visibility
Fig. 4: Experimental setup. (a) Motion-capture room; the inset shows one of the motion-capture cameras. (b) The five objects: block, mouse, mug, mustard bottle, and phone.
Fig. 5: Object-pose errors over the 40 motion-capture sessions (top: translation; bottom: symmetry-aware rotation). (a) Median error of each tracker binned by the visibility vdGT ; shading marks the occluded range. (b) PCF on the 24,285 occluded frames as a function of the error threshold; dotted lines mark the 20 mm and 15 ∘ thresholds used for ST20 and SR15 , and the numbers give the AUC. Medians are used in (a) because diverging baseline errors would dominate a mean.
Fig. 6: Qualitative results with increasing hand occlusion from top to bottom. Columns: object, 3DGS mesh, and the input image overlaid with the ground-truth pose (green) and the pose of FoundationPose, ICG+, and VICON. Last column: rotated view of the tracked hand, visualized as a MANO [ 27 ] mesh, and the VICON object pose with the contact points and force arrows of the recognized FSRs, colored by the calibrated normal force.
Reconstructing the motion of objects from videos is a key component for embodied AI and robot manipulation. While diverse approaches to object pose tracking have been studied, they rely heavily on strong external priors, such as depth data or 3D templates, and remain highly vulnerable to severe occlusions by hand grasps despite the use of explicit masks. In this work, we present ComPose, a 6DoF object tracking framework designed for hand-aware object pose estimation from RGB video. Rather than treating the hand purely as an occluder, our method harmonizes hand motions as a \textit{complementary cue} for object tracking. In detail, we recover a variety of object motions over time by combining object and hand cues from foundation models within a unified tracking pipeline. Here, ComPose adaptively selects informative hand joints, combines object- and hand-derived cues for motion estimation, and refines the resulting object motion using visible geometric evidence and a learned correction. We further enforce the temporal consistency over both rotation and translation, yielding stable 3D object trajectories over time without any external smoothing. Extensive experiments show that our method is accurate, efficient, and robust under severe hand occlusion and geometric ambiguity. In addition, the resulting trajectories can also effectively transfer to downstream robot manipulation by enabling robots to reconstruct human actions from online videos.
Dexterous in-hand manipulation requires continuous 6D pose tracking, yet the manipulating fingers inevitably occlude the object from the camera. We study how to structure the sparse haptic signals already available on multi-fingered hands, including proprioception, proximal force/torque, and binary contact, to complement a pretrained visual pose tracker under occlusion. We propose a kinematic-aware finger-level encoder and systematically compare it against four alternative designs through three levels of evaluation: per-frame refinement, sequential open-loop tracking, and closed-loop manipulation. Our experiments reveal that (i) per-frame evaluation cannot distinguish encoder quality, while sequential tracking amplifies architectural differences by up to 15 times; (ii) the structured encoder learns task-specific cross-modal gating, using vision exclusively for translation and dedicating one attention head to haptics for rotation, without explicit supervision; and (iii) compact finger-level tokenization with 4 tokens outperforms both flat fusion and joint-level representations, which suppress vision through norm dominance. We validate that improved tracking yields higher success in a downstream reorientation task and provide qualitative real-world demonstrations. Our project page is available at https://cold-young.github.io/kine-fuse/.
Chanyoung Ahn, Jaesung Lee, Sungwoo Park +1
Center for Humanoid Research, KIST, Seoul, 02792 South Korea · Korea University, Seoul, 02841 South Korea
Perceiving physical contact is fundamental to dexterous manipulation. While robots often rely on dedicated hardware tactile sensors, humans exhibit a remarkable ability to infer contact by integrating visual information with an innate sense of their body's pose and movement. Inspired by this embodied perceptual skill, we investigate whether a robot can learn to infer contact from vision, an approach that also offers a scalable alternative to tactile hardware specifically for binary contact estimation, which faces practical challenges in cost, fragility, and integration. We present NoContactNoWorries, a transformer-based multimodal framework that fuses RGB-D vision with the robot's proprioception to infer binary contact states as a pseudo-tactile signal for hand-object interactions. We validate by training a single contact prediction model on multiple objects and show that the inferred contact signal supports downstream reinforcement learning agents for in-hand object reorientation, generalizing to novel objects. Experiments in both simulation and on a real-world robot validate our approach, highlighting the feasibility of inferring contact from vision and proprioception. Project Page: https://soham2560.github.io/no-contact-no-worries/
Soham Patil, Avirup Das, Sourabh Bhosale +1
Robotics Research Center (RRC), International Institute of Information Technology (IIIT), Hyderabad, India · Department of Computer Science, The University of Manchester