DRHeC: Differentiable Rendering for Hand-Eye Calibration with RGB-Based Gradients
Authors: Xiaotian Zhang, Yusheng Wang, Naoya Kagawa, Noritaka Takamura, Keiji Okuhara, Hiroyasu Baba, Jun Ota
Organizations: Department of Precision Engineering, School of Engineering, The University of Tokyo, Tokyo, Japan · Research into Artifacts, Center for Engineering (RACE), School of Engineering, The University of Tokyo, Tokyo, Japan · FA Products Business Unit, DENSO WAVE INCORPORATED, Aichi, Japan
Accurate hand-eye calibration is crucial for precision manipulation. Traditional methods rely on markers, with their precision dependent on marker accuracy and observability. In contrast, markerless methods, such as learning-based approaches, use deep neural networks to directly extract keypoints or features from images, enabling the computation of hand-eye transformation with a single image and without the need for physical markers. Recently, differentiable rendering-based methods for hand-eye calibration have leveraged physical models to render binary masks and compare them with observations, enabling hand-eye calibration without fiducial markers in the calibration stage and providing interpretable optimization. While the state-of-the-art differentiable rendering methods achieve remarkable accuracy, the use of binary masks can result in the loss of internal profile details, reducing precision. Additionally, these methods can also suffer from unstable optimization and local minima. In this study, we propose a novel RGB-based differentiable rendering framework that provides richer geometric and appearance cues by incorporating color and mask geometric features, thereby improving calibration accuracy and optimization stability. Additionally, we propose a mask-guided image-to-image translation method to ensure explicit preservation of color and geometric consistency throughout the translation. Our approach is validated through both simulation and real-world experiments, with results demonstrating strong accuracy and robustness and clear improvements over existing differentiable rendering methods. Our method achieves a grasping success rate of 88.9% and insertion success rate of 57.4% on the UR5e real-world experiment, outperforming the state-of-the-art differentiable rendering hand-eye calibration method EasyHeC by 46.3 and 48.1 percentage points, respectively.
Figures & tables
Fig. 1: Comparison of optimization processes using EasyHeC [ 25 ] (first row), IPE [ 24 ] (second row), and DRHeC (third row) on the synthetic UR5e dataset. All methods start from the same initial pose and target ground-truth pose. EasyHeC and IPE struggle to converge when the initial pose deviates significantly from the ground truth due to large or incorrect gradients, often leading to optimization failure. In those methods, the manipulator flies out of the frame during the optimization process. In contrast, the proposed DRHeC method effectively mitigates these issues and successfully converges.
Fig. 2: The architecture of the differentiable renderer hand-eye calibration. The observed real image is first transformed into a simulation-like (sim-like) image using I2IT, which is then compared with the rendered image for differentiable renderer hand-eye calibration.
Fig. 3: Distributions of calibration errors under varying levels of initial perturbations for different differentiable rendering methods: EasyHeC [ 25 ] , IPE [ 24 ] , RGB, RGBC (RGB and Mask Centroid), and DRHeC. (a) Low perturbations, (b) Moderate perturbations, (c) High perturbations.
Perturbation Level
Method
Translation Error (meters)
Rotation Error (radians)
Success Rate (%)
Q1
Q3
Max
Mean
Q1
Q3
Max
Mean
Low (LP)
EasyHeC [ 25 ]
0.0002
0.0005
1.1282
0.0089
0.0010
0.0029
0.7012
0.0122
91.5
IPE [ 24 ]
0.0008
0.0021
0.0084
0.0016
0.0034
0.0100
0.0693
0.0095
100.0
RGB
0.0002
0.1037
1.8703
0.3419
0.0008
0.4508
1.7163
0.3254
79.0
RGB+Mask Centroid
0.0001
0.0003
0.0012
0.0003
0.0007
0.0016
0.0080
0.0013
100.0
DRHeC
0.0002
0.0003
0.0013
0.0003
0.0006
0.0015
0.0053
0.0012
100.0
TABLE I: Experimental Results on the Synthetic UR5e Dataset.
Method
20px
30px
40px
50px
100px
150px
200px
Dream [ 23 ]
0.16
0.23
0.29
0.33
0.52
0.62
0.64
OK [ 22 ]
0.34
0.54
0.66
0.69
0.88
0.93
0.95
IPE (box) [ 24 ]
-
-
-
0.65
0.94
0.95
0.95
IPE (cylinder) [ 24 ]
-
-
-
0.80
0.91
0.93
0.95
IPE (CAD) [ 24 ]
-
-
-
0.74
0.90
0.95
0.95
EasyHeC [ 25 ]
0.35
0.55
0.75
0.90
0.95
0.95
1.00
TABLE II: Evaluation results of 2D PCK on the real-world dataset.
Fig. 4: Experimental setup for two robot tasks. (a) Grasping experiment to verify the robot’s ability to grasp an object after hand-eye calibration. (b) Insertion experiment to verify the robot’s ability to accurately insert into a hole after hand-eye calibration.
Fig. 5: Example of real-world image translated into a simulation-like image for differentiable rendering-based hand-eye calibration.
Fig. 6: Visual comparison of real-world RGB images and predicted simulation-like images from different I2IT methods across robot poses. DRHeC shows better performance in color consistency for challenging robot configurations compared to other methods.
Method
MSE
MAE
IoU
PSNR (dB)
SSIM
CycleGAN [ 46 ]
294.35
3.19
0.9300
23.68
0.9725
CUT [ 47 ]
354.60
3.46
0.9284
22.76
0.9680
GeoMaskGAN [ 48 ]
304.74
3.26
0.9286
23.54
0.9705
DRHeC (ours)
292.12
3.19
0.9350
23.70
0.9722
TABLE III: Performance of I2IT methods in the real-world experiment.
Method
Marker-based
Differentiable-rendering-based (Markerless)
Tsai et al. [ 7 ]
Daniilidis et al. [ 8 ]
Single3D (closed-form) [ 2 ]
Single3D (iterations) [ 2 ]
Wang et al. [ 21 ]
EasyHeC [ 25 ]
DRHeC
Grasping
5.6%
9.3%
3.7%
96.3%
24.1%
42.6%
88.9%
Insertion
5.6%
9.3%
1.8%
81.5%
16.7%
9.3%
57.4%
TABLE IV: Success rates of different experiments on UR5e robot.
Fig. 7: Reprojection of calibration results on real-world images for two scenarios (1 and 2). Results from DRHeC are shown in the left column (1-a and 2-a), whereas those from EasyHeC are shown in the right column (1-b and 2-b).
Method
Marker-based
Differentiable-rendering-based ( Markerless )
Tsai et al. [ 7 ]
Daniilidis et al. [ 8 ]
Single3D (closed-form) [ 2 ]
Single3D (iterations) [ 2 ]
Wang et al. [ 21 ]
EasyHeC [ 25 ]
DRHeC
Grasping
11.1%
11.1%
38.9%
94.4%
44.4%
55.6%
77.8%
Insertion
11.1%
11.1%
27.8%
83.3%
22.2%
33.3%
77.8%
TABLE V: Success rates of different experiments on DENSO VS060 robot.
Fig. 8: Illustrative example of pose ambiguity in binary-mask-based calibration. RGB images (a-c) and corresponding binary masks (d-f) represent three distinct robot poses: (a,d): [0°, -90°, 0°, -90°, 0°, 0°]; (b,e): [90°, -90°, 0°, -90°, 0°, 0°]; and (c,f): [-90°, -90°, 0°, -90°, 0°, 0°]. Despite the 180° difference in Joint 1 between (b) and (c), their binary masks (e,f) appear nearly identical due to projection ambiguity. Methods relying on binary-mask loss may have difficulty distinguishing these poses during optimization. In contrast, DRHeC incorporates RGB information, which can help distinguish such cases.
Fig. 9: Ablation study on (a) λcentroid , (b) λarea , and (c) threshold for the two-step optimization parameters. Each pair shows translation and rotation errors, respectively.
λcentroid
0
0.1
0.01
0.001
0.0001
Translation
86.06
10.67
5.05
158.62
188.10
Rotation
0.00
0.04
0.02
0.03
0.00
TABLE VI: Ablation study results of the λcentroid over 2000 iterations.
λarea
0
0.1
0.01
0.001
0.0001
Translation
134.86
111.17
21.81
16.99
2.30
Rotation
0.05
0.11
0.01
0.04
0.01
TABLE VII: Ablation study results of the λarea over 2000 iterations.
Threshold
0
0.20
0.25
0.30
0.50
Translation (mm)
183.83
18.32
13.65
28.23
150.11
Rotation (rad)
0.02
0.02
0.01
0.03
0.07
TABLE VIII: Ablation study results of the threshold for the two-step optimization over 2000 iterations.
Method
MSE
MAE
IoU
PSNR (dB)
SSIM
Real world image
0.0099
0.0162
0.6111
20.88
0.9477
I2ITed img
0.0026
0.0053
0.8520
25.95
0.9727
TABLE IX: Ablation study results of the I2IT module.
This work presents an RGB-D imaging-based approach to marker-free hand-eye calibration using a novel implementation of the iterative closest point (ICP) algorithm with a robust point-to-plane (PTP) objective formulated on a Lie algebra. Its applicability is demonstrated through comprehensive experiments using three well known serial manipulators and two RGB-D cameras. With only three randomly chosen robot configurations, our approach achieves approximately 90% successful calibrations, demonstrating 2-3x higher convergence rates to the global optimum compared to both marker-based and marker-free baselines. We also report 2 orders of magnitude faster convergence time (0.8 +/- 0.4 s) for 9 robot configurations over other marker-free methods. Our method exhibits significantly improved accuracy (5 mm in task space) over classical approaches (7 mm in task space) whilst being marker-free. The benchmarking dataset and code are open sourced under Apache 2.0 License, and a ROS 2 integration with robot abstraction is provided to facilitate deployment.
Martin Huber, Huanyu Tian, Christopher E. Mower +4
School of Biomedical Engineering & Imaging Sciences, King’s College London, London, UK · Noah’s Ark Lab, Huawei, London, UK · independent contributor
Purpose: Marker-based tracking of surgical robots is occlusion-prone in cluttered operating rooms. We evaluate stereo differentiable rendering for marker-free, real-time robot pose tracking, potentially improving safety, reducing setup time, and enabling multi-robot interaction. Methods: We extend the markerless pose estimation framework roboreg to online dynamic tracking via (i) sequential optimisation that propagates pose estimates across frames with motion-adaptive hyperparameter tuning, and (ii) CUDA stream parallelisation of segmentation and optimisation, combined with CUDA-graph accelerated segmentation. We evaluate on 38 unobstructed and 5 occluded displacement sequences with static start/end ground-truth calibrations and dynamic marker-based reference tracking. Results: We achieve real-time 1080p tracking at 30 fps (up from 14 fps for vanilla roboreg), matching the camera frame rate. Accuracy reaches 1.7 cm / 0.6 deg against static ground truth and 1.2 cm mean 3D error over 27,460 frames against the marker-based reference (1.53 cm over 1,242 occluded frames). Our method outperforms FoundationPose by 11% in dynamic estimation (63% under occlusion) and 250% in static estimation, with 6x faster inference. Conclusions: Stereo differentiable rendering enables real-time, high-resolution marker-free surgical robot tracking, on par with marker-based approaches and surpassing foundation-model baselines.
Yanghe Hao, Martin Huber, Christos Bergeles +1
School of Biomedical Engineering & Imaging Sciences, King’s College London, London, United Kingdom.
This article proposes a general optimization framework for solving hand-eye calibration problem. Unlike traditional methods, an iterative algorithm based on Lie algebra that achieves approximately global optimal solutions is developed. During the optimization process, the method strictly preserves the structural constraints of the calibration parameters and enables synchronized updates between calibration parameters. Recognizing that data used in real-word hand-eye calibration often contain uncertainty, especially in over-loading and large workspace industrial robot scenarios, which can significantly degrade accuracy, and accurately modeling such uncertainty is inherently difficult, this article avoids explicit uncertainty modeling. Instead, an uncertainty metric to evaluate the relative uncertainty between data sources is introduced and used to dynamically refine the iterative process. To further enhance convergence efficiency, an effective initial solution generation method that improves overall stability and accuracy is designed. Numerical simulations and real-world experiments validate the effectiveness of the proposed approach, and in synthetic datasets, the proposed approach improves the estimation accuracy by at least 67% under high-uncertainty conditions compared with the existing methods.
Yanjia Chen, Xiangfei Li, Huan Zhao +4
State Key Laboratory of Intelligent Manufacturing Equipment and Technology and the School of Mechanical Science and Engineering, Huazhong University of Science and Technology, Wuhan, 430074, China.