Organizations: Department of Artificial Intelligence, School of Engineering, Westlake University · School of Artificial Intelligence, Shanghai Jiao Tong University
Learning dexterous humanoid loco-manipulation from human demonstrations requires transferring not only human motion, but also the coordinated interaction structure underlying the demonstrated behavior. This is challenging because embodiment differences distort the coupling among body motion, wrist placement, finger articulation, and object interaction, while kinematically accurate references may still be difficult to realize under robot dynamics. We present DexWeave, a unified framework that connects interaction-consistent motion retargeting with anatomy-aware whole-body policy learning. DexWeave first employs a two-stage retargeting procedure that initializes body and hand motions with specialized solvers and subsequently performs coupled refinement over the upper-body interaction chain while preserving lower-body support. The resulting references are tracked by an anatomy-aware Transformer policy that represents anatomical regions as structured tokens and uses directed masked attention to model their dependencies, with object information selectively conditioning the upper-body pathway for dexterous interaction. The policy jointly outputs body and dexterous-hand actions and is trained directly with reinforcement learning, without pretrained tracking policies, teacher-student distillation, or subsequent residual refinement. DexWeave improves retargeting fidelity and interaction consistency while achieving higher manipulation performance and faster policy convergence than MLP baselines. We further deploy the learned policies on a physical Unitree G1 humanoid equipped with Inspire dexterous hands, demonstrating dexterous whole-body loco-manipulation in the real world. See our project page (https://dexweave.github.io) for videos.
Figures & tables
Figure 2: Overview of DexWeave. Given a human body–hand–object demonstration, DexWeave first constructs an interaction-consistent robot reference through specialized body–hand initialization and coupled refinement of the upper-body interaction chain. An anatomy-aware whole-body loco-manipulation policy then realizes the reference under robot dynamics using regional anatomical tokens, directed masked attention, and selective object conditioning, producing executable dexterous humanoid loco-manipulation skills.
Figure 3: Anatomy-aware policy architecture. Left: Robot observations are factorized into anatomical tokens and coordinated by a masked whole-body Transformer. Object information is fused one-way into the waist, arm, and hand branches, while the original waist feature is retained for object-independent leg control. Seven group-specific action heads produce the full body-and-hand action. Right: Attention-mask visualization. Rows attend to columns. Light and dark green indicate allowed attention in the robot Transformer and object-fusion Transformer, respectively.
Penetration
Foot Skating
Contact Preservation
Dataset
Method
Duration ↓
Max Depth (cm) ↓
Duration ↓
Max Vel. (m/s) ↓
Duration ↑
Distance (cm) ↓
LAFAN1
OmniRetarget
0.125
2.934
0.091
0.416
N/A
N/A
GMR
0.183
4.314
0.010
0.548
N/A
N/A
SOMA
0.968
6.379
0.115
0.565
N/A
N/A
DexWeave (Ours)
∼0
2.575
0
0
N/A
N/A
OMOMO
OmniRetarget
0.131
2.735
∼0
1.455
0.644
10.289
Table 1: Whole-body feasibility and coarse interaction preservation on LAFAN1 and OMOMO. Penetration and foot skating measure feasibility; contact preservation is evaluated only on OMOMO. Durations are fractions, and ∼0 denotes a near-zero value. Yellow and underlined mark the best and second-best values within each dataset.
Penetration
Hand Alignment
Dataset
Method
Duration ↓
Max Depth (cm) ↓
Pri. Err. (mm) ↓
Sec. Err. (mm) ↓
Palm Err. (°) ↓
GRAB
OmniRetarget + DexPilot + IK
0.212
2.269
16.585
13.368
11.555
OmniRetarget + SBR + IK
0.190
2.362
14.925
14.659
9.530
DexWeave (Ours)
0.001
1.263
4.342
14.476
3.652
HUMOTO
OmniRetarget + DexPilot + IK
0.522
3.597
32.389
31.018
14.444
OmniRetarget + SBR + IK
0.511
3.582
29.999
28.436
12.476
Table 2: Fine-grained interaction retargeting on GRAB and HUMOTO. Penetration duration is the fraction of colliding frames. Primary and secondary errors measure thumb/index and remaining-finger positions; palm error measures normal alignment. Lower is better. Yellow and underlined mark the best and second-best values within each dataset. Full comparison can be found in Appendix D .
Figure 4: Qualitative retargeting results across interaction levels. Each row shows a sequence overview (left) and hand–object close-ups (right).
Figure 6
Figure 6: Sim-to-sim and sim-to-real transfer of dexterous loco-manipulation policies. Policies are trained in IsaacLab, evaluated in MuJoCo for sim-to-sim transfer, and deployed on a physical Unitree G1 with Inspire dexterous hands for sim-to-real transfer.
Appendix figures & tables10 assets
Supplementary material from the paper’s appendix.
Appendix
Term
Target
Weight or distance scale
Body and support tracking
Pelvis pose
World position; mapped x/y/z axes.
82k ; (2.4,2.4,3.2)k
Torso pose
World position; mapped axes.
11k ; (1.05,1.05,1.45)k
Shoulder / elbow positions
World-frame shoulders; torso-relative elbows.
4.2k ; 1.2k
Wrist pose
Interaction-dependent position; mapped axes.
240k ; (2.1,1.7,0.9)k
Foot positions
Support-aware sole center; heel and toe positions; sole-corner heights.
115k ; 48k each; 260k each
Appendix
Table 4: Body-optimization objectives. Position and orientation-axis terms use squared Euclidean and squared angular errors, respectively; posture terms act on joint coordinates. Weights are per landmark, axis, or joint, with k=103 . Foot and knee weights are listed before support/contact modulation. Lengths are in meters and angles in radians unless stated otherwise.
Term
Target and activation
Weight or setting
Codebook and retrieval
Driver sampling
Six mimic-feasible drivers: thumb yaw/pitch and four finger flexions; interval sampling followed by forward kinematics.
{0.15,0.50,0.85} ; NB=36=729 per hand
Finger descriptors
Wrist-to-tip, chain-root, and up to three segment unit vectors; accumulated bending and tip reach divided by chain length.
Up to 15+2 components per finger
Inter-finger distances
Selected tip-to-tip distances normalized by the mean wrist-to-tip radius.
6 pairs; at most 91 descriptor components overall
Masked retrieval
Observed-component distance and exponential driver blending in Eq. equation 13 .
K=12 ; τ=0.06
Geometric fitting
Appendix
Table 5: Hand-initialization settings. Fitting weights are per landmark or driver and multiply squared position or driver-coordinate errors. Position targets are wrist-local; lengths are in meters and angles in radians. Here k=103 .
Term
Target and activation
Weight or setting
Interaction alignment
Thumb / index fingertips
World-frame and wrist-local targets in Eq. equation 14 .
2.0 each
Other fingertips
Corresponding targets for the middle, ring, and little fingers.
0.20 each
Hand orientation
Palm normal, wrist-forward direction, and lateral axis.
0.55 , 0.42 , 0.12
Position and posture anchors
Wrist / palm / elbow positions
Soft position anchors; wrist and palm anchors relax with increasing interaction activation.
0.08 / 0.06 / 0.0125
Appendix
Table 6: Coupled-refinement objectives. Position weights are divided by ℓt,s2 and joint-coordinate weights by squared feasible ranges. Activation factors are applied as described in the text. Angular residuals use radians.
Observation
Dimension
Input details
Actor observations
Reference body and hand joint positions
41
Corresponding joint groups
Reference non-finger joint velocities
29
Legs, waist, arms
Reference-anchor position history relative to torso
7×3
Base
Reference-anchor orientation relative to torso
6
Base
Base angular velocity
3
Base
Appendix
Table 7: Actor and critic observations for dexterous loco-manipulation. Dimensions count scalar inputs before encoding. The last column specifies regional assignments for actor inputs and feature composition for critic inputs.
Term
Weight
Scale σ
Whole-body tracking
Global torso position / orientation
1.5/1.5
0.20/0.20
Aligned body position / orientation
1.0/1.0
0.30/0.40
Global body linear / angular velocity
1.0/1.0
1.0/3.14
Body joint positions
1.0
0.40
Hand and object interaction
Appendix
Table 8: Reward configuration for dexterous loco-manipulation. Tracking rewards use the listed weights and error scales σ , and penalty terms use direct coefficients. Lengths are in meters and angles in radians.
Setting
Value
Setting
Value
Parallel environments
4,096
Physics / control frequency
200 / 50 Hz
Steps per environment
24
Samples per update
98,304
Learning epochs / minibatches
5 / 4
Initial learning rate
10−3
PPO clipping parameter
0.2
Target KL divergence
0.01
Discount γ / GAE λ
0.99 / 0.95
Value-loss coefficient
1.0
Maximum gradient norm
1.0
Initial action standard deviation
1.0
Appendix
Table 9: PPO and simulation settings.
Quantity
Model
Range or value
Contact and inertial properties
Static friction
Coefficient
[0.4,1.8]
Dynamic friction
Coefficient
[0.35,1.5]
Restitution
Fixed coefficient
0
Foot contact offset
Distance parameter
[5,20] mm
Body mass
Multiplicative
[0.9,1.1]
Appendix
Table 10: Domain randomization and measurement settings. Randomized quantities are sampled uniformly over the listed ranges; ±a denotes [−a,a] . Multiplicative factors are relative to nominal values. Fixed settings are marked explicitly.
Penetration
Hand Alignment
Dataset
Method
Duration ↓
Max Depth (cm) ↓
Pri. Err. (mm) ↓
Sec. Err. (mm) ↓
Palm Err. (°) ↓
GRAB
OmniRetarget + DexPilot
0.183
2.123
209.353
207.745
52.122
OmniRetarget + DexPilot + IK
0.212
2.269
16.585
13.368
11.555
OmniRetarget + SBR
0.184
2.120
210.524
212.985
52.122
OmniRetarget + SBR + IK
0.190
2.362
14.925
14.659
9.530
DexWeave (Ours)
0.001
1.263
4.342
14.476
3.652
Appendix
Table 11: Complete fine-grained interaction retargeting results on GRAB and HUMOTO. Penetration duration is the fraction of colliding frames. Primary and secondary errors measure thumb/index and remaining-finger positions; palm error measures normal alignment. Lower is better. Yellow and underlined mark the best and second-best values within each dataset.
Penetration
Hand Alignment
Dataset
Stage
Duration ↓
Max Depth (cm) ↓
Pri. Err. (mm) ↓
Sec. Err. (mm) ↓
Palm Err. (deg) ↓
GRAB
Decoupled
0.157
1.778
36.033
36.995
0.189
Refined
0.001
1.263
4.342
14.476
3.652
HUMOTO
Decoupled
0.204
2.568
32.703
26.090
0.159
Refined
0.007
2.854
7.198
11.702
4.439
Appendix
Table 12: Retargeting-stage ablation. Decoupled is the body-and-hand initialization; refined is the final output.
Attention setting
Complete Ratio (%) ↑
Body pos. (cm) ↓
Object pos. (cm) ↓
Object rot. (deg) ↓
MR dense
91.30
4.79
3.62
8.23
MO bidirectional
99.71
4.63
3.88
9.19
MR and MO (DexWeave)
100
4.58
3.81
8.47
Appendix
Table 13: Policy-mask ablation. Results are averaged over two reference motions using checkpoints saved every 100 iterations during the final 1,000 training iterations, with five evaluation seeds per checkpoint and reference. Dense MR and bidirectional MO each relax one attention mask.
Learning humanoid-object interaction requires coordinating whole-body balance, locomotion, and dexterous hand contact to control both robot and object motion. Human demonstrations provide examples of coordinated interaction, but transferring these behaviors to humanoid robots requires learning how to establish and maintain effective contacts under different embodiments and dynamics. We present Weave, a unified framework for learning whole-body dexterous humanoid-object interaction from captured human demonstrations. Weave first converts captured human-object interactions into executable robot-object references through contact-aware retargeting and approach-motion completion. At its core is a contact- and geometry-aware policy that jointly commands 29 body joints and 12 actuated finger joints across multiple objects and interaction sequences. Evaluation across nine objects yields a 92.5% success rate on trained interactions and, without any additional training, 65.0% on sequences never seen during training. We additionally release ~9,000 physically executed rollouts spanning ~23 hours, providing robot-object trajectories with contact annotations for downstream interaction-policy learning and physically consistent HOI motion generation. Project website: https://xiaohu-art.github.io/Weave/
Liu Cao, Xingze Wu, Jingzhi Cui +4
1Tsinghua IIIS · 3Dalian University of Technology · 4The Chinese University of Hong Kong +1
Humanoid loco-manipulation is often simplified into a stop-and-go process: walking to an object, stopping to manipulate it, and then resuming locomotion. It also commonly relies on low degree-of-freedom (DoF) end effectors that behave like an open-close grasp primitive. We introduce CoorDex, a learning pipeline that converts high-dimensional body and dexterous hand control into coordinated latent residual control, enabling high-DoF dexterous loco-manipulation on the move. Starting from simulated whole-body and hand demonstrations, CoorDex trains privileged motion tracking teachers for the humanoid body and dexterous hand, distills them into proprioception-conditioned latent priors, and uses the frozen priors as the action space for downstream residual reinforcement learning. A coordinated latent residual policy composes these priors through shared task context and separate body-hand residual heads, preserving natural whole-body motion while improving finger-level contact reliability. CoorDex enables a Unitree G1 humanoid with a 20-DoF WUJI hand to execute dexterous manipulation while in motion, including non-stop bottle grasping and carrying, fridge door opening on the move, and cube pick-and-turn. Ablations on the walk-grasp-carry task show that joint-space PPO, joint-space hand control, and monolithic latent prediction all fail under the same reward budget, while the latent-prior interface and coordinated residual structure make high-dimensional contact-rich loco-manipulation trainable. Project Page: https://skevinci.github.io/coordex/
Sikai Li, Shuning Li, Zhenyu Wei +3
University of North Carolina at Chapel Hill · University of California, Berkeley
Recent work in humanoid whole-body control has found success with a simple recipe: retarget human motion to robot kinematic references, then train policies via reinforcement learning (RL) to track them. But how does this recipe transfer to dexterous manipulation? The answer is not obvious, as manipulation involves complex, contact-rich dynamics and requires delicate regulation of contact modes and forces. We present REGRIND, a minimalist retargeting-guided RL pipeline that learns dexterous manipulation policies from a single human demonstration. REGRIND retargets human hand-object motion to a robot reference that preserves hand-object spatial and contact relationships, trains a residual RL policy in simulation to track object-centric keypoints along that reference, and transfers the resulting policy zero-shot to hardware with careful system identification. The resulting policies produce fluid, human-like behavior on two different multi-fingered hands across contact-rich tool-use tasks, including operating a pair of scissors and turning a screwdriver. Through systematic hardware experiments, we identify and analyze the key factors that govern sim-to-real transfer in dexterous manipulation, offering practical guidance for retargeting-based learning in contact-rich settings. Videos and code are available at https://yunhaifeng.com/REGRIND.