Torque observations in reinforcement learning remain challenging because simulated and measured torque differ in scale, offset, and noise. In this paper, we propose a simple torque observation alignment method for robots with direct-drive (DD) actuators, in which motor current maps linearly to joint torque through a motor-type-specific torque constant K_tau. First, dynamometer calibration identifies K_tau* and corrects the scale mismatch between simulated and real torque. Second, the method uses delta_tau(t) = tau(t) - tau(t-1) as the observation in both domains to eliminate the constant offset instead of using the direct torque tau(t), which carries a domain-dependent bias. Third, Gaussian noise obtained from the dynamometer measurement data is injected during the learning process. To validate the proposed method, we train a teacher-student grasping policy entirely in simulation and deploy the distilled student on a multifingered DD gripper. The deployed policy performs proprioceptive grasping using only joint positions and torque differences. We conduct an ablation study comparing the proposed method with alternative alignment variants on nine in-distribution (ID) objects. The proposed method achieves 100% grasp success. These results demonstrate that the proposed alignment method improves the robustness of zero-shot policy transfer on the DD gripper against real-world torque-observation mismatches.
Fig. 3: Direct-drive gripper platform. (a) DD actuator-based multifingered gripper. (b) Direct-Drive (DD) actuators integrated into a single finger. (c) Joint and link configuration.
Fig. 4: Dynamometer-based identification of the effective torque constants. (a) Experimental setup consisting of a test motor, a load motor, and an inline torque sensor. (b) Measured shaft torque versus quadrature current for the GL40 motor. (c) The corresponding result for the GL60 motor. Dashed lines denote the linear regression fits.
Motor
Datasheet Kτ
Calibrated Kτ∗
R2
RMSE [Nm]
[Nm/A]
[Nm/A]
GL60 KV25
0.4500
0.3260
0.9940
0.0098
GL40 KV70
0.1500
0.1392
0.9992
0.0023
TABLE I: Datasheet and dynamometer-identified torque constants with regression accuracy.
Motor
Gaussian model ϵm,t
Injected std. σm [Nm]
GL60 KV25
N(0,9.688465×10−5)
0.009843
GL40 KV70
N(0,5.313025×10−6)
0.002305
TABLE II: Motor-type-specific Gaussian noise derived from dynamometer regression residuals and injected into simulated torque observations.
Fig. 5: Teacher-student policy training and deployment pipeline.
Teacher (PPO)
Parameter
Value
Learning rate
10−3
Steps per environment
24
PPO update epochs
5
Minibatches
6
Activation
ELU
TABLE III: Training hyperparameters.
Shape
Dimensions [cm]
A
B
C
Cuboid
Base length × height
6×6
8×8
10×8
Cylinder
Radius × height
4×4
4.5×4.5
5×5
Sphere
Radius
4
4.5
5
TABLE IV: Dimensions of the training objects.
Parameter
Range
Robot: Initial MCP Joint Offset (rad)
+U[−0.06,0.06]
Robot: Initial PIP Joint Offset (rad)
+U[−0.35,0.35]
Object: Initial Position ( x , y ) (m)
+U[−0.01,0.01]
Object: Initial Yaw (rad)
U[−π,π]
Object: Mass
×U[0.5,3.0]
Actuator: P (Stiffness) Gain
×U[0.85,1.05]
TABLE V: Randomization applied during teacher training.
Fig. 7: Objects used in the real-world evaluation. The evaluation includes 21 objects: nine in-distribution (ID) objects and twelve out-of-distribution (OOD) objects. The ID objects belong to the cuboid, cylinder, and sphere categories used during simulation training. The OOD objects are real-world items with previously unseen shapes, materials, and surface properties that fit within the gripper workspace.
Policy
Observation
Noise
Success rate
Simulation
Real-world
Position Only
[qt]
N/A
68.9%
15.6%
Scale + Noise
[qt,τt]
On
100.0%
26.7%
Scale + Bias
[qt,Δτt]
Off
100.0%
62.2%
Our Method
[qt,Δτt]
On
100.0%
100.0%
TABLE VI: Simulation and real-world success rates of student policies πs .
ID Objects
OOD Objects
#
Object
Rate
#
Object
Rate
1
Cuboid A
100%
10
Box
100%
2
Cuboid B
100%
11
Stuffed toy
100%
3
Cuboid C
100%
12
Tennis ball
90%
4
Cylinder A
100%
13
Plastic container
100%
5
Cylinder B
100%
14
Wine glass
100%
TABLE VII: Real-world grasping success rates on ID and OOD objects.
Fig. 8: Representative real-world grasping sequences and motor-current-derived Δτ trajectories for the complete method and two Δτ ablations on the DD gripper. The sequences show grasp outcomes alongside the torque-change observations. (a) Our Method , using calibrated Kτ∗ and noise injection, completes the grasp–lift–hold task. (b) Scale + Bias , using calibrated Kτ∗ without noise injection, exhibits a grasping/lifting failure. (c) Bias + Noise , using the datasheet Kτ with noise injection, exhibits a holding failure. In the plot legends, F1–F3 denote the three fingers, while MCP, PIP, and DIP denote the three joints of each finger.
Human-like dexterous hands with multiple fingers offer human-level manipulation capabilities but remain difficult to train the control policies that can deploy on real hardware due to contact-rich physics and imperfect actuation. We present a sim-to-real reinforcement learning method that leverages dense tactile feedback combined with joint torque sensing to explicitly regulate physical interactions. To enable effective sim-to-real transfer, we introduce (i) a computationally fast tactile simulation that computes distances between dense virtual tactile units and the object via parallel forward kinematics, providing high-rate, high-resolution touch signals needed by RL; (ii) a current-to-torque calibration that eliminates the need for torque sensors on dexterous hands by mapping motor current to joint torque; and (iii) actuator dynamics modeling with randomization to account for non-ideal torque-speed effects and bridge the actuation gaps. Using an asymmetric actor-critic PPO pipeline, we train policies entirely in simulation and deploy them directly to a five-finger hand. The resulting policies demonstrate two essential human-hand skills: (1) command-based controllable grasp force tracking and (2) reorientation of objects in the hand, both of which are robustly executed without fine-tuning on the robot. By combining tactile and torque in the observation space with scalable sensing and actuation modeling, our system provides a practical solution to achieve reliable dexterous manipulation. To our knowledge, this is the first demonstration of controllable grasping on a multi-finger dexterous hand trained entirely in simulation and transferred zero-shot on real hardware.
Zhe Zhao, Zhibin Li, Yilin Ou +1
State Key Laboratory of Networking and Switching Technology Beijing University of Posts and Telecommunications, China · University College London, United Kingdom
A policy tuned for one robot often behaves differently on another, whether due to the sim-to-real gap, unknown payloads, or the differing dynamics of two instances of the same robot. In contact-rich, dynamic manipulation, even small motion discrepancies can result in failure to track reference motion, since they disrupt the timing and modes of contact. Common remedies, such as domain randomization or system identification, either produce overly conservative task policies or require data that must be recollected for each robot or payload. We introduce the Torque Adaptation Module (TAM), a learned module that adapts the torque commands sent to the robot to match the behavior of an ideal robot. TAM operates between the low-level controller that tracks the policy's actions and the robot's torque interface. It includes a history encoder that embeds proprioceptive history into a latent state and a torque adaptor that computes residual torque corrections. Because TAM depends only on proprioceptive history and not on policy observations, or the action space, the same TAM weights can be reused to adapt policies with different action spaces (joint targets, end-effector targets, or direct torques). The policies themselves do not need to be trained with domain randomization of robot parameters. Instead, we offload the need for domain randomization to TAM by training it entirely in randomized simulation, using multi-robot pretraining followed by a robot-specific fine-tuning step that still requires no real-robot data. We evaluate TAM zero-shot on a real Franka Panda robot across dynamic manipulation tasks that include a vision-based box pushing policy (from RL), a flip policy (from BC), and an MPC ball-on-plate balancing. Our experiments show that TAM improves zero-shot real-robot execution compared to online system identification and RMA baselines and enables robust dynamic manipulation performance.
Dongwon Son, Florian Shkurti, Jason Lee +3
1KAIST · *Work done during an internship at the Allen Institute for AI. · 2Allen Institute for AI +2
Sim-to-real transfer in robot learning is often limited by discrepancies between the ideal actuator dynamics assumed during policy training and the nonlinear, hardware-dependent behavior of physical motors. While conventional approaches attempt to bridge this gap by increasing simulator fidelity through system identification, domain randomization, or learned actuator models, we introduce an alternative paradigm: actuator reality shaping. Instead of modifying the simulator to match the real world, our method shapes the closed-loop behavior of physical actuators to match the idealized second-order reference dynamics used in simulation. By equipping each joint with a two-degree-of-freedom feedforward--feedback controller, we decouple reference-response shaping from robust stabilization, thereby providing a standardized actuator interface for reinforcement learning policies. As a result, policies trained only with the prescribed reference model can be deployed zero-shot on real hardware without task-level fine-tuning or learned actuator models. We validate the approach on a single-joint high-gear-ratio servo under external loads and a 7-DOF robotic arm reaching task, where actuator reality shaping substantially reduces sim-to-real tracking error and improves zero-shot task performance compared with standard servo-control and representative real-to-sim-to-real baselines. We further demonstrate zero-shot transfer on a wheeled-legged robot driving over a slope and a humanoid robot walking, suggesting that actuator reality shaping can serve as a reusable interface for robot learning across diverse hardware platforms. Project page: https://syamamori.github.io/ActuatorRealityShaping.github.io/
Satoshi Yamamori, Koji Ishihara, Kenjiro Minamikawa +4
Graduate School of Informatics, Kyoto University, Kyoto, Japan · Dept. of Brain Robot Interface, Computational Neuroscience Labs, ATR, Kyoto, Japan