Torque observations in reinforcement learning remain challenging because simulated and measured torque differ in scale, offset, and noise. In this paper, we propose a simple torque observation alignment method for robots with direct-drive (DD) actuators, in which motor current maps linearly to joint torque through a motor-type-specific torque constant K_tau. First, dynamometer calibration identifies K_tau* and corrects the scale mismatch between simulated and real torque. Second, the method uses delta_tau(t) = tau(t) - tau(t-1) as the observation in both domains to eliminate the constant offset instead of using the direct torque tau(t), which carries a domain-dependent bias. Third, Gaussian noise obtained from the dynamometer measurement data is injected during the learning process. To validate the proposed method, we train a teacher-student grasping policy entirely in simulation and deploy the distilled student on a multifingered DD gripper. The deployed policy performs proprioceptive grasping using only joint positions and torque differences. We conduct an ablation study comparing the proposed method with alternative alignment variants on nine in-distribution (ID) objects. The proposed method achieves 100% grasp success. These results demonstrate that the proposed alignment method improves the robustness of zero-shot policy transfer on the DD gripper against real-world torque-observation mismatches.
Fig. 3: Direct-drive gripper platform. (a) DD actuator-based multifingered gripper. (b) Direct-Drive (DD) actuators integrated into a single finger. (c) Joint and link configuration.
Fig. 4: Dynamometer-based identification of the effective torque constants. (a) Experimental setup consisting of a test motor, a load motor, and an inline torque sensor. (b) Measured shaft torque versus quadrature current for the GL40 motor. (c) The corresponding result for the GL60 motor. Dashed lines denote the linear regression fits.
Motor
Datasheet Kτ
Calibrated Kτ∗
R2
RMSE [Nm]
[Nm/A]
[Nm/A]
GL60 KV25
0.4500
0.3260
0.9940
0.0098
GL40 KV70
0.1500
0.1392
0.9992
0.0023
TABLE I: Datasheet and dynamometer-identified torque constants with regression accuracy.
Motor
Gaussian model ϵm,t
Injected std. σm [Nm]
GL60 KV25
N(0,9.688465×10−5)
0.009843
GL40 KV70
N(0,5.313025×10−6)
0.002305
TABLE II: Motor-type-specific Gaussian noise derived from dynamometer regression residuals and injected into simulated torque observations.
Fig. 5: Teacher-student policy training and deployment pipeline.
Teacher (PPO)
Parameter
Value
Learning rate
10−3
Steps per environment
24
PPO update epochs
5
Minibatches
6
Activation
ELU
TABLE III: Training hyperparameters.
Shape
Dimensions [cm]
A
B
C
Cuboid
Base length × height
6×6
8×8
10×8
Cylinder
Radius × height
4×4
4.5×4.5
5×5
Sphere
Radius
4
4.5
5
TABLE IV: Dimensions of the training objects.
Parameter
Range
Robot: Initial MCP Joint Offset (rad)
+U[−0.06,0.06]
Robot: Initial PIP Joint Offset (rad)
+U[−0.35,0.35]
Object: Initial Position ( x , y ) (m)
+U[−0.01,0.01]
Object: Initial Yaw (rad)
U[−π,π]
Object: Mass
×U[0.5,3.0]
Actuator: P (Stiffness) Gain
×U[0.85,1.05]
TABLE V: Randomization applied during teacher training.
Fig. 7: Objects used in the real-world evaluation. The evaluation includes 21 objects: nine in-distribution (ID) objects and twelve out-of-distribution (OOD) objects. The ID objects belong to the cuboid, cylinder, and sphere categories used during simulation training. The OOD objects are real-world items with previously unseen shapes, materials, and surface properties that fit within the gripper workspace.
Policy
Observation
Noise
Success rate
Simulation
Real-world
Position Only
[qt]
N/A
68.9%
15.6%
Scale + Noise
[qt,τt]
On
100.0%
26.7%
Scale + Bias
[qt,Δτt]
Off
100.0%
62.2%
Our Method
[qt,Δτt]
On
100.0%
100.0%
TABLE VI: Simulation and real-world success rates of student policies πs .
ID Objects
OOD Objects
#
Object
Rate
#
Object
Rate
1
Cuboid A
100%
10
Box
100%
2
Cuboid B
100%
11
Stuffed toy
100%
3
Cuboid C
100%
12
Tennis ball
90%
4
Cylinder A
100%
13
Plastic container
100%
5
Cylinder B
100%
14
Wine glass
100%
TABLE VII: Real-world grasping success rates on ID and OOD objects.
Fig. 8: Representative real-world grasping sequences and motor-current-derived Δτ trajectories for the complete method and two Δτ ablations on the DD gripper. The sequences show grasp outcomes alongside the torque-change observations. (a) Our Method , using calibrated Kτ∗ and noise injection, completes the grasp–lift–hold task. (b) Scale + Bias , using calibrated Kτ∗ without noise injection, exhibits a grasping/lifting failure. (c) Bias + Noise , using the datasheet Kτ with noise injection, exhibits a holding failure. In the plot legends, F1–F3 denote the three fingers, while MCP, PIP, and DIP denote the three joints of each finger.
State Key Laboratory of Networking and Switching Technology Beijing University of Posts and Telecommunications, China · University College London, United Kingdom