HACo: Learning Haptic Active Compliance for Force-Aware Dexterous Manipulation
Authors: Naisheng Ye, Yinzhe Zhou, Junkai Zhao, Yuhang Lu, Checheng Yu, Zhenjie Yang, Pengwei Wang, Hongyang Li
Organizations: The University of Hong Kong, Hong Kong SAR, China · Beijing Academy of Artificial Intelligence (BAAI), Beijing, China · Johns Hopkins University, Baltimore, MD, USA
Contact-rich dexterous manipulation requires policies that translate physical feedback into motion commands while regulating interaction loads across evolving multi-contact interactions. This requires haptic observations of contact state and action supervision showing how commands should adapt. Existing policies often overlook complementary fingertip tactile and joint-torque feedback, while common action targets either encode excessive loading or omit motion constrained by the object. We introduce HACo, a Haptic Active Compliance policy that learns force-regulating actions directly from haptic feedback. Compliance-regulated teleoperation converts operator inputs into controller-executable compliant actions that preserve motion intent while regulating loads. HACo learns these actions directly, using command-state discrepancy as auxiliary compliant-intent supervision. It combines local fingertip tactile responses with joint-torque feedback capturing load transmission through the articulated hand, including contacts beyond tactile coverage. A Compliance Grounding Module uses gated haptic cross-attention to ground action generation in the evolving haptic state, enabling closed-loop force regulation without explicit online contact modeling. We evaluate HACo on a real-world benchmark covering multi-contact friction, tangential interaction, fragile curved-surface contact, rotational torque, and deformable-object manipulation. Across 20 trials per task, HACo achieves an 83% mean success rate, compared with 35% for the strongest evaluated baseline. These results demonstrate active compliance across diverse force-sensitive dexterous manipulation tasks.
Figures & tables
Fig. 2: Overview of HACo. The haptic expert aligns fingertip wrench and deformation features with joint-torque features from the same finger before cross-finger fusion. The Compliance Grounding Module (right) lets action features query this haptic representation through gated haptic cross-attention alongside vision-language conditioning. Trained with conditional flow matching, the action expert learns compliant actions from force-regulated demonstrations, with auxiliary compliant-intent supervision capturing the discrepancy between compliant hand commands and observed hand states.
Fig. 3: Compliance-regulated teleoperation. Motion retargeting provides nominal arm and hand references. Cartesian admittance and hand force regulation convert them into controller-executable compliant references while synchronized robot state and haptic observations are recorded.
Fig. 4: Hand action semantics under contact. Nominal retargeting can command excessive contact force, while observed configurations alone omit motion constrained by contact. The compliant command preserves task-directed motion with force regulation. Its discrepancy from the observed state, Δqtci , provides compliant-intent supervision; δqt∗ denotes the correction to the nominal command. Overlaid configurations illustrate these distinctions.
Fig. 5: Teleoperation and experimental platform. The operator uses MANUS Metagloves Pro and VIVE Trackers to capture finger articulation and wrist pose, which are retargeted to the Sharpa hands and UR5 end-effectors. A ZED Mini and two wrist-mounted D405 cameras provide visual observations; robot state, tactile feedback, and joint torques are recorded synchronously.
Method
Insert poker cards
Open book
Draw on balloon
Unscrew cap
Squeeze toothpaste
Mean
GR00T [ 38 ]
3/20
0/20
1/20
7/20
4/20
15%
GR00T + Tactile
5/20
2/20
1/20
9/20
5/20
22%
ViTacFormer [ 1 ]
0/20
1/20
0/20
2/20
1/20
4%
T-Rex [ 8 ]
4/20
6/20
2/20
12/20
11/20
35%
HACo
18/20
17/20
14/20
19/20
15/20
83%
TABLE I: Policy comparison on the real-world force-sensitive dexterous manipulation benchmark. Entries report successful rollouts out of 20; Mean is the macro-average success rate across five tasks.
Configuration
Insert poker cards
Open book
Draw on balloon
Unscrew cap
Squeeze toothpaste
Mean
HACo
18/20
17/20
14/20
19/20
15/20
83%
Haptic perception
w/o Haptic Feedback
7/20
2/20
3/20
8/20
7/20
27% (-56%)
w/o Tactile Feedback
8/20
5/20
7/20
13/20
12/20
45% (-38%)
w/o Torque Feedback
15/20
16/20
11/20
14/20
12/20
68% (-15%)
w/o Coupled Encoding
17/20
14/20
12/20
13/20
14/20
70% (-13%)
TABLE II: Ablation results. Entries report successful rollouts out of 20; parenthesized values denote the decrease in mean success rate from full HACo.
Fig. 6: Representative HACo rollouts on the real-world dexterous force benchmark. Each row shows seven temporally ordered frames of one execution, progressing from contact establishment to task completion. Better viewed in videos.
Compliance is essential for dexterous manipulation, yet existing solutions often rely on external tactile or force sensors that are costly, fragile, and difficult to deploy on low-cost robot hands. We propose a proprioception-driven framework that learns contact-aware compliance cues from motor current and joint states. Since motor current is closely related to actuator torque, it provides an intrinsic signal for perceiving contact force, object resistance, and grasp stability without additional sensing hardware. Rather than estimating external wrenches or commanding torque, our method predicts a compliance reference position: an ideal joint-position target for a standard PD controller whose induced position error generates appropriate grasping force. This position-based formulation is compatible with mainstream teleoperation and policy-learning pipelines, while enabling the robot to adapt interaction forces from real-time proprioceptive feedback. Thus, motor current serves not only as a force proxy but also as a learnable proprioceptive contact signal for compliance reference prediction. Experiments on multiple dexterous hands and contact-rich tasks, including fragile object handling, sustained surface contact, thin-object retrieval, and dynamic load adaptation, show stable compliant grasping, safer and more efficient teleoperation, and improved downstream policy learning without external tactile or force sensors.
Dexterous grasping depends on contact regulation, not motion alone. Stable manipulation requires fingers to maintain appropriate object loading as contacts slip, deform, or become visually occluded. Existing cross-embodiment dexterous policies unify motion through retargeted hand poses or latent actions, but force feedback remains tied to each hand's sensing and actuation, limiting transfer. This work introduces a cross-embodiment force-position interface for contact-aware manipulation across heterogeneous dexterous hands. Motion intent is represented in a shared hand-pose latent, while each hand's effort signal is calibrated through system identification into physical joint torque in N.m. These torques are mapped to fingertip forces and compact per-finger load descriptors, giving the policy comparable observations of where the hand should move and how the object is loaded. Using this interface, a flow-matching visuomotor policy is trained on vision, proprioception, and calibrated contact, with structured visual masking that encourages reliance on force under grasp-relevant occlusion. The same calibrated signal drives a hybrid force-position controller for demonstration collection and execution, keeping force targets consistent across training and deployment. Experiments across structurally different hands show that calibrated contact feedback enables transferable compliant grasping, with learned primitives reusable in long-horizon manipulation pipelines.
Soofiyan Atar, Yao-Ting Huang, Michael Yip
Department of Electrical and Computer Engineering University of California San Diego, United States
Contact-rich manipulation requires robots to regulate both motion and interaction forces, yet achieving adaptive compliance remains a fundamental challenge. Learning from real-world data is costly and risky, while simulation-based approaches struggle with the sim-to-real gap in contact dynamics; existing sim-to-real methods either require real-world adaptation or sacrifice adaptive compliance by relying on isotropic compliant controllers. Our key insight is that force regulation decomposes into a time-varying but simulation-transferable directional component and a dynamics-sensitive but manually tunable magnitude component. We instantiate this directional component as two policy outputs, a task frame and a control mode vector, predicted by a visuomotor policy adapted from a pre-trained VLA model and trained via imitation learning on automatically generated simulation demonstrations. At deployment, an admittance controller integrates these predictions with human-specified stiffness and target wrench values to realize adaptive compliance. Our approach achieves adaptive compliance using only simulation data and can benefit from large-scale VLA pre-training. Extensive real-world experiments on four contact-rich tasks, microwave opening, peg-in-hole insertion, whiteboard wiping, and door opening, demonstrate strong task success rates and robustness to external disturbances. Project page: https://yifei-y.github.io/project-pages/TDC/.
Yifei Yang, Anzhe Chen, Zhenjie Zhu +6
Zhejiang University · Zhejiang Humanoid Robot Innovation Center