Dexterous manipulation policies trained in simulation often fail to transfer to the real world because of errors in contact timing and force regulation. Yet these policies can retain useful multi-finger coordination for task progression. We propose ReDex, a framework for adapting a simulation-trained base policy to the real world by correcting local contact failures and incorporating tactile feedback. Starting from a proprioception-only base policy, ReDex allows a human operator to physically correct contact failures at selected fingers under compliant control during real-world rollouts, while the frozen base policy continues to control the remaining fingers. These rollouts combine base policy execution, human-corrected finger motion, and fingertip force observations. We reconstruct force-informed targets from these rollouts to train a standalone force-conditioned policy via behavior cloning. This design reduces human correction effort, enables learning of contact regulation from real-world interaction, and introduces force feedback into a proprioception-only policy without tactile simulation or complex full-hand teleoperation. We evaluate ReDex on two challenging, contact-rich dexterous manipulation tasks on real hardware. Compared with sim-to-real transferred base policies, ReDex increases Object Flipping success rate from 14% to 86% across two objects and average Screwdriver Rotation progress from 26.0% to 95.3% across three objects.
Figures & tables
Fig. 1: ReDex overview. ReDex repairs local sim-to-real failures through finger-level physical correction and reconstruction of training targets. We show (a) Screwdriver Rotation and (b) cellphone Object Flipping. Top: real-world contact failures during base policy execution. Bottom: physical correction of selected fingers. RL denotes the base policy. The corrected rollouts are relabeled with force-informed targets and distilled into a standalone force-conditioned policy via behavior cloning (BC).
Fig. 2: ReDex pipeline. Top: ReDex follows a three-stage pipeline: (1) train a proprioception-only dexterous policy in simulation, (2) collect real-world finger-level physical corrections through a compliant framework while the base policy controls the remaining fingers, and (3) train a standalone force-conditioned policy from the reconstructed real-world rollouts. Bottom: We detail the correction and supervision-reconstruction process: (2a) finger-level shared control enables selective physical guidance without interrupting the remaining hand motion; (2b) measured fingertip forces are mapped to force-informed joint targets that encode the desired contact load; (2c) these targets are smoothly blended with the evolving base policy targets at intervention onset and release; and (2d) the reconstructed targets form the training supervision for the final policy, which uses proprioception and fingertip force for autonomous deployment.
Fig. 3: Qualitative rollout. Each column shows four frames of autonomous execution after learning from human corrections, ordered from top to bottom. From left to right, we show Object Flipping with a cuboid and a cellphone, and Screwdriver Rotation with normal (N), small (S), and large (L) screwdrivers. Insets visualize changes in screwdriver orientation. T denotes the elapsed time, in seconds, from the start of the corresponding rollout.
Screwdriver Rotation (progress. %)
Object Flipping (successes / trials)
Method
Normal
Small
Large
Overall
Cuboid
iPhone
Overall
RMA [ 2 ]
26
33
19
26.0
4/25
3/25
7/50
Gated Corr. [ 7 ]
42
—
—
—
3/25
—
—
Continuous Corr. [ 8 ]
36
—
—
—
1/25
—
—
ReDex (Ours)
95
99
92
95.3
20/25
23/25
43/50
TABLE I: Real-robot performance over 25 autonomous trials per object–method pair. Prog. is mean discretized rotation progress, quantized in 90∘ increments and reported as a percentage; SR is the number of successful trials. Screwdriver Overall is the mean across the three object variants, while Flipping Overall pools trials across the two object variants. Overall results are reported only when all corresponding object variants were evaluated.
Variant
Force obs.
Virtual target
Transition blending
SR.
Motion only
–
–
–
4/10
w Tactile obs.
✓
–
–
6/10
w/o virtual target
✓
–
✓
7/10
w/o blending
✓
✓
–
6/10
Ours
✓
✓
✓
9/10
TABLE II: Component ablation on real-robot Object Flipping (Cuboid). All variants use motion supervision. Checkmarks denote enabled components; SR gives successes over 10 matched trials.
Tactile Obs.
Encoder
Flip SR.
Taxels
MLP
6/10
Taxels
GRU
2/10
Net force
GRU
8/10
Net force
MLP
9/10
TABLE III: Ablation of tactile observation, representation, and encoder on real-robot Object Flipping (Cuboid). All variants follow a matched 10-trial protocol. Net force stacks one fingertip triaxial resultant per fingertip.
Fig. 4: Human correction across objects. Bars show the share of recorded rollout time with at least one finger under human correction. Values inside the bars give human / total rollout time in minutes, rounded to 0.1 min. The five circles represent the thumb, index, middle, ring, and little fingers from left to right; blue marks guided fingers and gray marks fingers controlled by the base policy.
Full kinesthetic teaching
Ours
Guided fingers
Thumb, index
Thumb
Replay-validated trajectories
30
30
Recorded rollout duration (min)
13.2
9.9
Human time (min)
13.2
2.4
Mean duration per trajectory (s)
26.4
19.8
TABLE IV: Human guidance effort on Object Flipping (cuboid). Both methods yield 30 trajectories that successfully replay without human intervention. Reported rollout durations sum the recorded execution time of these trajectories; they exclude failed attempts, resets, and replay validation. Both methods use force-informed joint position targets for replay.
Human-like dexterous hands with multiple fingers offer human-level manipulation capabilities but remain difficult to train the control policies that can deploy on real hardware due to contact-rich physics and imperfect actuation. We present a sim-to-real reinforcement learning method that leverages dense tactile feedback combined with joint torque sensing to explicitly regulate physical interactions. To enable effective sim-to-real transfer, we introduce (i) a computationally fast tactile simulation that computes distances between dense virtual tactile units and the object via parallel forward kinematics, providing high-rate, high-resolution touch signals needed by RL; (ii) a current-to-torque calibration that eliminates the need for torque sensors on dexterous hands by mapping motor current to joint torque; and (iii) actuator dynamics modeling with randomization to account for non-ideal torque-speed effects and bridge the actuation gaps. Using an asymmetric actor-critic PPO pipeline, we train policies entirely in simulation and deploy them directly to a five-finger hand. The resulting policies demonstrate two essential human-hand skills: (1) command-based controllable grasp force tracking and (2) reorientation of objects in the hand, both of which are robustly executed without fine-tuning on the robot. By combining tactile and torque in the observation space with scalable sensing and actuation modeling, our system provides a practical solution to achieve reliable dexterous manipulation. To our knowledge, this is the first demonstration of controllable grasping on a multi-finger dexterous hand trained entirely in simulation and transferred zero-shot on real hardware.
Zhe Zhao, Zhibin Li, Yilin Ou +1
State Key Laboratory of Networking and Switching Technology Beijing University of Posts and Telecommunications, China · University College London, United Kingdom
Sim-to-real transfer remains a critical bottleneck for deploying dexterous manipulation policies learned in simulation to real-world robots. Existing approaches rely on manually designed domain randomization or task-specific adaptation, limiting their generalizability across diverse manipulation scenarios. We present DexSim2Real, an integrated framework that leverages vision-language foundation models to bridge the sim-to-real gap for dexterous manipulation. Our system combines three components: (1) Foundation Model-Guided Domain Randomization (FM-DR), which uses a vision-language model as a visual realism critic to optimize simulation parameters via closed-loop CMA-ES, complementing text-based approaches like DrEureka with direct visual feedback; (2) a Tactile-Visual Cross-Attention Policy (TVCAP) that adapts cross-attention visuo-tactile fusion to zero-shot sim-to-real RL; and (3) a Progressive Skill Curriculum (PSC) that builds on LLM-based task decomposition with a difficulty scheduler tailored to contact-rich dexterous tasks. Extensive experiments on six challenging manipulation tasks with blinded evaluation demonstrate that DexSim2Real achieves a 78.2% average real-world success rate, outperforming DrEureka and DeXtreme while reducing the sim-to-real performance gap to only 8.3%.
Zijian Zeng, Fei Ding, Huiming Yang +2
Tsinghua University · Alibaba Group · Bengbu University +1
Human videos are an abundant source of dexterous manipulation behaviors, but they lack tactile information that is crucial for contact-rich interaction. This raises a fundamental question: can robots learn deployable visual-tactile dexterous manipulation policies from human video demonstrations without robot-side data collection? We present DEX-X, a framework for learning visual-tactile dexterous manipulation from human videos through simulation. Our key insight is that simulation can serve as a tactile completion engine. Given monocular human demonstrations, DEX-X reconstructs hand-object interactions in simulation, where physically grounded contact dynamics provide tactile supervision unavailable in the original videos. Leveraging this recovered tactile information, we train visual-tactile dexterous manipulation policies and distill them into deployable policies operating on point-cloud observations and tactile sensing. We demonstrate zero-shot sim-to-real transfer on a dexterous hand-arm platform across diverse grasping and contact-rich tool-use tasks. The teacher policy achieves 65.9% average success across six task categories in simulation, while the distilled visual-tactile policy achieves 93% success on real-world cube picking and 53% on the challenging table-cleaning task. Zero-shot generalization to unseen object geometries is also observed on object-picking tasks. Our results suggest that simulated interaction is a key bridge between human videos and deployable dexterous manipulation policies, providing the missing physical supervision needed for scalable robot skill learning from Internet-scale human video data.
Ruoqu Chen, Feixiang Ruan, Liu Cao +9
Tsinghua University · Shanghai Qizhi Institute · Sharpa +2