DRHeC: Differentiable Rendering for Hand-Eye Calibration with RGB-Based Gradients
Authors: Xiaotian Zhang, Yusheng Wang, Naoya Kagawa, Noritaka Takamura, Keiji Okuhara, Hiroyasu Baba, Jun Ota
Organizations: Department of Precision Engineering, School of Engineering, The University of Tokyo, Tokyo, Japan · Research into Artifacts, Center for Engineering (RACE), School of Engineering, The University of Tokyo, Tokyo, Japan · FA Products Business Unit, DENSO WAVE INCORPORATED, Aichi, Japan
Accurate hand-eye calibration is crucial for precision manipulation. Traditional methods rely on markers, with their precision dependent on marker accuracy and observability. In contrast, markerless methods, such as learning-based approaches, use deep neural networks to directly extract keypoints or features from images, enabling the computation of hand-eye transformation with a single image and without the need for physical markers. Recently, differentiable rendering-based methods for hand-eye calibration have leveraged physical models to render binary masks and compare them with observations, enabling hand-eye calibration without fiducial markers in the calibration stage and providing interpretable optimization. While the state-of-the-art differentiable rendering methods achieve remarkable accuracy, the use of binary masks can result in the loss of internal profile details, reducing precision. Additionally, these methods can also suffer from unstable optimization and local minima. In this study, we propose a novel RGB-based differentiable rendering framework that provides richer geometric and appearance cues by incorporating color and mask geometric features, thereby improving calibration accuracy and optimization stability. Additionally, we propose a mask-guided image-to-image translation method to ensure explicit preservation of color and geometric consistency throughout the translation. Our approach is validated through both simulation and real-world experiments, with results demonstrating strong accuracy and robustness and clear improvements over existing differentiable rendering methods. Our method achieves a grasping success rate of 88.9% and insertion success rate of 57.4% on the UR5e real-world experiment, outperforming the state-of-the-art differentiable rendering hand-eye calibration method EasyHeC by 46.3 and 48.1 percentage points, respectively.
Figures & tables
Fig. 1: Comparison of optimization processes using EasyHeC [ 25 ] (first row), IPE [ 24 ] (second row), and DRHeC (third row) on the synthetic UR5e dataset. All methods start from the same initial pose and target ground-truth pose. EasyHeC and IPE struggle to converge when the initial pose deviates significantly from the ground truth due to large or incorrect gradients, often leading to optimization failure. In those methods, the manipulator flies out of the frame during the optimization process. In contrast, the proposed DRHeC method effectively mitigates these issues and successfully converges.
Fig. 2: The architecture of the differentiable renderer hand-eye calibration. The observed real image is first transformed into a simulation-like (sim-like) image using I2IT, which is then compared with the rendered image for differentiable renderer hand-eye calibration.
Fig. 3: Distributions of calibration errors under varying levels of initial perturbations for different differentiable rendering methods: EasyHeC [ 25 ] , IPE [ 24 ] , RGB, RGBC (RGB and Mask Centroid), and DRHeC. (a) Low perturbations, (b) Moderate perturbations, (c) High perturbations.
Perturbation Level
Method
Translation Error (meters)
Rotation Error (radians)
Success Rate (%)
Q1
Q3
Max
Mean
Q1
Q3
Max
Mean
Low (LP)
EasyHeC [ 25 ]
0.0002
0.0005
1.1282
0.0089
0.0010
0.0029
0.7012
0.0122
91.5
IPE [ 24 ]
0.0008
0.0021
0.0084
0.0016
0.0034
0.0100
0.0693
0.0095
100.0
RGB
0.0002
0.1037
1.8703
0.3419
0.0008
0.4508
1.7163
0.3254
79.0
RGB+Mask Centroid
0.0001
0.0003
0.0012
0.0003
0.0007
0.0016
0.0080
0.0013
100.0
DRHeC
0.0002
0.0003
0.0013
0.0003
0.0006
0.0015
0.0053
0.0012
100.0
TABLE I: Experimental Results on the Synthetic UR5e Dataset.
Method
20px
30px
40px
50px
100px
150px
200px
Dream [ 23 ]
0.16
0.23
0.29
0.33
0.52
0.62
0.64
OK [ 22 ]
0.34
0.54
0.66
0.69
0.88
0.93
0.95
IPE (box) [ 24 ]
-
-
-
0.65
0.94
0.95
0.95
IPE (cylinder) [ 24 ]
-
-
-
0.80
0.91
0.93
0.95
IPE (CAD) [ 24 ]
-
-
-
0.74
0.90
0.95
0.95
EasyHeC [ 25 ]
0.35
0.55
0.75
0.90
0.95
0.95
1.00
TABLE II: Evaluation results of 2D PCK on the real-world dataset.
Fig. 4: Experimental setup for two robot tasks. (a) Grasping experiment to verify the robot’s ability to grasp an object after hand-eye calibration. (b) Insertion experiment to verify the robot’s ability to accurately insert into a hole after hand-eye calibration.
Fig. 5: Example of real-world image translated into a simulation-like image for differentiable rendering-based hand-eye calibration.
Fig. 6: Visual comparison of real-world RGB images and predicted simulation-like images from different I2IT methods across robot poses. DRHeC shows better performance in color consistency for challenging robot configurations compared to other methods.
Method
MSE
MAE
IoU
PSNR (dB)
SSIM
CycleGAN [ 46 ]
294.35
3.19
0.9300
23.68
0.9725
CUT [ 47 ]
354.60
3.46
0.9284
22.76
0.9680
GeoMaskGAN [ 48 ]
304.74
3.26
0.9286
23.54
0.9705
DRHeC (ours)
292.12
3.19
0.9350
23.70
0.9722
TABLE III: Performance of I2IT methods in the real-world experiment.
Method
Marker-based
Differentiable-rendering-based (Markerless)
Tsai et al. [ 7 ]
Daniilidis et al. [ 8 ]
Single3D (closed-form) [ 2 ]
Single3D (iterations) [ 2 ]
Wang et al. [ 21 ]
EasyHeC [ 25 ]
DRHeC
Grasping
5.6%
9.3%
3.7%
96.3%
24.1%
42.6%
88.9%
Insertion
5.6%
9.3%
1.8%
81.5%
16.7%
9.3%
57.4%
TABLE IV: Success rates of different experiments on UR5e robot.
Fig. 7: Reprojection of calibration results on real-world images for two scenarios (1 and 2). Results from DRHeC are shown in the left column (1-a and 2-a), whereas those from EasyHeC are shown in the right column (1-b and 2-b).
Method
Marker-based
Differentiable-rendering-based ( Markerless )
Tsai et al. [ 7 ]
Daniilidis et al. [ 8 ]
Single3D (closed-form) [ 2 ]
Single3D (iterations) [ 2 ]
Wang et al. [ 21 ]
EasyHeC [ 25 ]
DRHeC
Grasping
11.1%
11.1%
38.9%
94.4%
44.4%
55.6%
77.8%
Insertion
11.1%
11.1%
27.8%
83.3%
22.2%
33.3%
77.8%
TABLE V: Success rates of different experiments on DENSO VS060 robot.
Fig. 8: Illustrative example of pose ambiguity in binary-mask-based calibration. RGB images (a-c) and corresponding binary masks (d-f) represent three distinct robot poses: (a,d): [0°, -90°, 0°, -90°, 0°, 0°]; (b,e): [90°, -90°, 0°, -90°, 0°, 0°]; and (c,f): [-90°, -90°, 0°, -90°, 0°, 0°]. Despite the 180° difference in Joint 1 between (b) and (c), their binary masks (e,f) appear nearly identical due to projection ambiguity. Methods relying on binary-mask loss may have difficulty distinguishing these poses during optimization. In contrast, DRHeC incorporates RGB information, which can help distinguish such cases.
Fig. 9: Ablation study on (a) λcentroid , (b) λarea , and (c) threshold for the two-step optimization parameters. Each pair shows translation and rotation errors, respectively.
λcentroid
0
0.1
0.01
0.001
0.0001
Translation
86.06
10.67
5.05
158.62
188.10
Rotation
0.00
0.04
0.02
0.03
0.00
TABLE VI: Ablation study results of the λcentroid over 2000 iterations.
λarea
0
0.1
0.01
0.001
0.0001
Translation
134.86
111.17
21.81
16.99
2.30
Rotation
0.05
0.11
0.01
0.04
0.01
TABLE VII: Ablation study results of the λarea over 2000 iterations.
Threshold
0
0.20
0.25
0.30
0.50
Translation (mm)
183.83
18.32
13.65
28.23
150.11
Rotation (rad)
0.02
0.02
0.01
0.03
0.07
TABLE VIII: Ablation study results of the threshold for the two-step optimization over 2000 iterations.
Method
MSE
MAE
IoU
PSNR (dB)
SSIM
Real world image
0.0099
0.0162
0.6111
20.88
0.9477
I2ITed img
0.0026
0.0053
0.8520
25.95
0.9727
TABLE IX: Ablation study results of the I2IT module.
State Key Laboratory of Intelligent Manufacturing Equipment and Technology and the School of Mechanical Science and Engineering, Huazhong University of Science and Technology, Wuhan, 430074, China.