FoLD: Force-Informed Learning for Dexterous Articulated Object Manipulation
Authors: Haowei Shen, Tingai Li, Yumeng Liu, Wenyuan Guang, Xuanze Yang, Qing Fang, Kai Xu, Ligang Liu, +1 more
Organizations: University of Science and Technology of China · Institute of AI for Industries, Chinese Academy of Sciences · Jiangsu Key Laboratory of AI for Industries · Shenzhen University
Transferring human demonstrations to dexterous robots remains challenging because differences in hand morphology and contact dynamics often cause retargeted motions to fail at producing the intended object behavior. We present \textbf{FoLD}, a framework for learning dexterous manipulation of articulated objects through explicit force guidance. FoLD compute compensatory force fields from human demonstrations together with the robot's current interaction state, yielding a force prior that promotes the demonstrated object motion. This force prior informs a residual policy that adapts retargeted hand motions to the contact requirements of the task. We evaluate FoLD on a public benchmark for articulated object manipulation, where it consistently outperforms state-of-the-art baselines across tasks and embodiments. We further validate FoLD on real dexterous robot platforms, demonstrating successful transfer of human manipulation skills to robot execution. Here is the link of our project page: https://gghgghgghgg.github.io/FoLD-project-page/.
Figures & tables
Fig. 2: Overview of FoLD. The imitator provides reference robot actions from a human demonstration. As the robot executes these actions, FoLD estimates a compensatory force field from the resulting object-motion deviations and the current contact map. The field is encoded into a latent representation that informs a residual policy, which corrects the reference actions to bring the object motion closer to the demonstration.
Fig. 3: Qualitative results on mixer, ketchup-bottle, and notebook manipulation (top to bottom). From left to right: human demonstrations, DexMachina, hand-only imitation, and FoLD, with two snapshots per group. Hand-only imitation uses only our pretrained imitator’s hand actions, with the object’s root pose and articulation fixed to the demonstrated state at each frame.
Hand
Method
SR ↑
Contact Mean ↑
Pos. (cm) ↓
Rot. ( ∘ ) ↓
Joint ( ∘ ) ↓
Inspire
DexMachina
0.511
0.154
52.27
42.65
34.26
Ours
0.700
0.248
18.19
23.28
20.20
[0.5pt/5pt] Allegro
DexMachina
0.596
0.043
20.07
49.38
22.51
Ours
0.750
0.061
23.62
40.51
12.18
[0.5pt/5pt] XHand
DexMachina
0.724
0.113
39.55
48.34
24.88
Ours
0.697
0.116
31.86
43.76
20.24
TABLE I: Comparison of DexMachina and FoLD across robotic hands. Both methods use the same 27 paired hand–task evaluations (aggregation details in Appendix D-B ). Best results for each hand are shown in bold.
Fig. 4: Bimanual manipulation across hand embodiments. Human configurations appear on the left, followed by robot configurations learned with FoLD.
Fig. 5: Additional FoLD manipulation examples. Each group presents three snapshots showing coordinated support and manipulation of articulated object parts.
Method
SR ↑
CWS Mean ↑
Pos. (cm) ↓
Rot. ( ∘ ) ↓
Joint ( ∘ ) ↓
CHORD
0.392±0.315
0.351±0.242
32.90±29.47
43.58±38.03
21.92±23.45
CHORD + FoLD
0.397±0.314
0.344±0.225
20.19±17.92
33.60±30.40
22.17±24.00
TABLE II: Comparison of CHORD and CHORD + FoLD on Sharpa-Wave dexterous hand. Values are mean ± sample standard deviation.
Fig. 6: Human demonstrations, CHORD, and CHORD + FoLD on box, notebook, and laptop manipulation (top to bottom). These snapshots illustrate hand configurations relative to the articulated object parts.
Residual policy
VOC curriculum
Force-field input
SR ↑
Contact Mean ↑
Pos. (cm) ↓
Rot. ( ∘ ) ↓
Joint ( ∘ ) ↓
✓
✗
✗
0.848
0.300
1.04
11.33
11.99
✓
✓
✗
1.000
0.275
1.34
5.55
13.30
✓
✓
✓
1.000
0.351
0.84
7.69
13.76
TABLE III: Component ablation on Inspire. Checkmarks and crosses indicate enabled and disabled components, respectively. All rows use the same frozen imitator; the last row is FoLD as evaluated in the main comparison. Bold indicates the best result.
Fig. 7: Real-robot demonstrations of laptop closing (upper two rows) and headphone folding (lower two rows). Each task pairs the human demonstration with open-loop replay of the FoLD-generated trajectory on two FR3 arms with Inspire hands.
Appendix figures & tables5 assets
Supplementary material from the paper’s appendix.
Appendix
Parameter
Value
Position proportional gain kp
200
Position derivative gain kd
20
Rotation proportional gain kR
20
Rotation derivative gain kω
2
Force command scale
25N
Torque command scale
3Nm
Appendix
TABLE IV: Online compensatory force-field construction parameters.
Parameter
Value
Parallel environments
4096
Rollout horizon
16
Learning rate
3×10−4
PPO mini-epochs
5
Minibatch size
32768
Discount factor γ
0.99
Appendix
TABLE V: PPO and simulation configuration used in the DexMachina and FoLD experiments.
Term
DexMachina
FoLD
Object tracking
1.0
1.0
Hand-pose imitation
0.3
0.3
Frame-to-frame motion consistency
0.0
0.1
Joint-reference tracking
0.3
0.3
Position-based contact tracking
3.0
3.0
High-contact-force penalty
0.1
0.1
Appendix
TABLE VI: Reward coefficients for the DexMachina baseline and FoLD stage-2 residual-policy training.
Task
Training identifier
ARCTIC sequence
Frame range
Length
Ketchup-100
ketchup-30-130-s01-use_01
s01/use_01
[30,130)
100
Box-200
box-30-230-s01-use_01
s01/use_01
[30,230)
200
Mixer-170
mixer-30-200-s01-use_01
s01/use_01
[30,200)
170
Ketchup-300
ketchup-30-330-s01-use_01
s01/use_01
[30,330)
300
Mixer-300
mixer-30-330-s01-use_01
s01/use_01
[30,330)
300
Notebook-300
notebook-30-330-s02-use_01
s02/use_01
[30,330)
300
Appendix
TABLE VII: ARCTIC demonstrations used in the four-hand DexMachina comparison. Frame ranges use half-open intervals.
Sequence
Source frames
Eval. states
Metric samples
s01_box_use_01
889
1184
1183
s07_box_grab_01
725
965
964
s01_notebook_use_01
669
890
889
s01_laptop_use_01
653
869
868
s01_waffleiron_use_01
605
805
804
Appendix
TABLE VIII: Complete sequences used in the CHORD comparison.
Recent work in humanoid whole-body control has found success with a simple recipe: retarget human motion to robot kinematic references, then train policies via reinforcement learning (RL) to track them. But how does this recipe transfer to dexterous manipulation? The answer is not obvious, as manipulation involves complex, contact-rich dynamics and requires delicate regulation of contact modes and forces. We present REGRIND, a minimalist retargeting-guided RL pipeline that learns dexterous manipulation policies from a single human demonstration. REGRIND retargets human hand-object motion to a robot reference that preserves hand-object spatial and contact relationships, trains a residual RL policy in simulation to track object-centric keypoints along that reference, and transfers the resulting policy zero-shot to hardware with careful system identification. The resulting policies produce fluid, human-like behavior on two different multi-fingered hands across contact-rich tool-use tasks, including operating a pair of scissors and turning a screwdriver. Through systematic hardware experiments, we identify and analyze the key factors that govern sim-to-real transfer in dexterous manipulation, offering practical guidance for retargeting-based learning in contact-rich settings. Videos and code are available at https://yunhaifeng.com/REGRIND.
Dexterous robot manipulation can benefit from the abundance of human demonstrations, but transferring such demonstrations to robot policies remains challenging. We present Contact Wrench Guidance from Human Demonstration in Robotic Dexterous Manipulation (CHORD), a framework for long-horizon manipulation of rigid and articulated objects with reinforcement learning. The key idea is object-centric contact wrench space guidance: we represent human and robot motions by the forces and torques they can induce on the object, enabling similarity to be measured by the induced instantaneous motions. This guidance makes reinforcement learning more scalable for contact-rich dexterous manipulation. We further introduce a large-scale simulation benchmark with 4,739 bimanual dexterous manipulation tasks, constructed from motion-capture datasets and reconstructed in-house videos. Evaluated on 1,831 benchmark tasks, CHORD achieves an average success rate of 82.12%, demonstrating strong scalability. CHORD also generalizes to whole-body manipulation from hand-only and third-person demonstrations, achieving a 90.77% success rate, and the learned policies transfer to the real world in both open-loop and closed-loop settings.
Learning humanoid-object interaction requires coordinating whole-body balance, locomotion, and dexterous hand contact to control both robot and object motion. Human demonstrations provide examples of coordinated interaction, but transferring these behaviors to humanoid robots requires learning how to establish and maintain effective contacts under different embodiments and dynamics. We present Weave, a unified framework for learning whole-body dexterous humanoid-object interaction from captured human demonstrations. Weave first converts captured human-object interactions into executable robot-object references through contact-aware retargeting and approach-motion completion. At its core is a contact- and geometry-aware policy that jointly commands 29 body joints and 12 actuated finger joints across multiple objects and interaction sequences. Evaluation across nine objects yields a 92.5% success rate on trained interactions and, without any additional training, 65.0% on sequences never seen during training. We additionally release ~9,000 physically executed rollouts spanning ~23 hours, providing robot-object trajectories with contact annotations for downstream interaction-policy learning and physically consistent HOI motion generation. Project website: https://xiaohu-art.github.io/Weave/
Liu Cao, Xingze Wu, Jingzhi Cui +4
1Tsinghua IIIS · 3Dalian University of Technology · 4The Chinese University of Hong Kong +1