Efficient and robust tracking of surgical robotic instruments is important for robot-assisted minimally invasive surgery, yet remains challenging due to the complexity of surgical scenes and the unconventional geometry of surgical instruments. Keypoint-based approaches are efficient, but their performance depends on reliable feature detection. Improving these detectors with real-world supervision is difficult because accurate real-world annotations are costly to obtain at scale. To address this limitation, we introduce a tracker-guided self-training framework that adapts a model pretrained on synthetic images to unlabeled real-world videos. Given measured robot joint states, an uncertainty-aware EKF recursively corrects the instrument pose and the observable joint angles by comparing projected model features with detected keypoints, shaft boundaries, and mask-derived cues. An RTS smoother subsequently refines the resulting trajectory, which is projected into pseudo-labels for fine-tuning the feature detector without laborious pose annotations. Experiments on real-world videos demonstrate consistent improvements from self-training across all evaluated keypoint metrics, and the resulting model outperforms prior approaches in both accuracy and runtime. The code and data will be released upon publication.
Figures & tables
Fig. 1: Self-training improves sim-to-real feature detection. Pretrained and self-trained models are compared on unseen real-world images with diverse orientations, distractions, and occlusions. Circles and dashed white lines denote manually labeled ground truth; plus signs and solid yellow lines denote model predictions. A tracker-guided self-training framework enables a synthetically pretrained model to bridge the gap to challenging, real-world scenes without tedious human labeling.
Fig. 2: Overview of the proposed self-training framework. A query-based multitask model predicts keypoint heatmaps, shaft edge heatmaps, and instrument masks from RGB video, which are decoded into geometric observations with associated uncertainty. Sequential multi-cue fusion combines these observations with robot kinematics to obtain smoothed pose trajectories, using an initial base-to-camera transform Tbc estimated by differentiable rendering [ 8 ] . An EKF propagates the instrument state between frames and updates it using the current frame’s observations, while a backward RTS smoother refines these pose estimates to produce a smoothed pose trajectory. Pseudo-label generation projects and renders the smoothed instrument configurations into feature labels for sim-to-real adaptation.
Fig. 3: Synthetic data generation. Isaac Sim generates labeled images with ground-truth keypoints and shaft edge annotations.
Metric
Pretrained
Ours
Improvement
Per-keypoint accuracy (PCK@20)
Gripper 1
91.59
93.70
+2.11 pp
Gripper 2
94.92
95.99
+1.07 pp
End effector
94.13
98.12
+3.99 pp
Pitch
82.76
91.94
+9.18 pp
Pitch side
94.60
99.90
+5.30 pp
TABLE I: Keypoint Detection Accuracy. Pretrained and self-trained detectors are evaluated on 16 SurgPose trajectories, with errors measured at the original 1400 × 986 resolution and PCK reported as percentages. The best results are shown in bold.
Model
Gripper (px) ↓
All Kpts. (px) ↓
Edge ↑
Unified detector [ 7 ]
22.24
–
0.9606
Pretrained
17.19
12.57
0.9886
Ours
14.90
11.38
0.9887
TABLE II: Real-World Feature Detection Accuracy. Pretrained and self-trained detectors are evaluated on the annotated real-world data, with mean keypoint errors reported in pixels and edge detection measured by the modified EA score [ 7 ] . The best and second-best results are shown in bold and underlined, respectively.
Observations
Mean error ↓ (px)
PCK@20 ↑ (%)
Kpts.
Shaft
Mask
EKF
RTS
EKF
RTS
✓
–
–
11.50
10.68
91.63
93.99
✓
✓
–
11.08
10.34
92.91
95.78
✓
✓
✓
10.05
9.22
94.36
96.58
TABLE III: Ablation of Multi-Cue Fusion. Shaft and mask cues are incrementally incorporated into pseudo-label generation using the pretrained detector on the training split. Mean keypoint error and PCK@20 are reported for both EKF and RTS. The best and second-best results are shown in bold and underlined, respectively.
Fig. 4: Qualitative tracking comparison. Representative results across different ex vivo backgrounds for a rendering-based tracker [ 10 ] and the proposed EKF tracker with observations from the pretrained and self-trained detectors. The estimated poses are rendered as green silhouettes, while the reference poses are highlighted by white skeletons. The self-trained detector provides more accurate alignment across challenging tissue appearances and unseen backgrounds compared to the rendering-based baseline and the pretrained model.
Method
Mean err. ↓
P95 err. ↓
PCK@20 ↑
FPS ↑
CMA-ES + KF [ 10 ]
20.53
60.47
69.17
28.78
Pretrained (w/ mask)
9.40
22.15
93.13
81.14
Pretrained (w/o mask)
9.77
22.32
92.96
178.64
Ours (w/ mask)
9.27
20.60
94.44
81.08
Ours (w/o mask)
9.23
20.33
94.72
179.07
TABLE IV: Tracking Accuracy and Speed. Results are reported over the 16 SurgPose evaluation trajectories. The best and second-best results are shown in bold and underlined, respectively.
Accurate and efficient tracking of surgical instruments is fundamental for Robot-Assisted Minimally Invasive Surgery. Although vision-based robot pose estimation has enabled markerless calibration without tedious physical setups, reliable tool tracking for surgical robots still remains challenging due to partial visibility and specialized articulation design of surgical instruments. Previous works in the field are usually prone to unreliable feature detections under degraded visual quality and data scarcity, whereas rendering-based methods often struggle with computational costs and suboptimal convergence. In this work, we incorporate CMA-ES, an evolutionary optimization strategy, into a versatile tracking pipeline that jointly estimates surgical instrument pose and joint configurations. Using batch rendering to efficiently evaluate multiple pose candidates in parallel, the method significantly reduces inference time and improves convergence robustness. The proposed framework further generalizes to joint angle-free and bi-manual tracking settings, making it suitable for both vision feedback control and online surgery video calibration. Extensive experiments on synthetic and real-world datasets demonstrate that the proposed method significantly outperforms prior approaches in both accuracy and runtime. Source code and data are available at https://github.com/hanyang-hu/online_dvrk_tracking.
Hanyang Hu, Zekai Liang, Florian Richter +1
Department of Electrical and Computer Engineering, University of California San Diego, La Jolla, CA 92093 USA.
Objective assessment of robotic surgery uses instrument kinematics, which must be reconstructed when only video is available. We introduce a kinematic reconstruction network for estimating instrument position, orientation and jaw angle from monocular video. Our visual representation combines global attention pooling of frozen DINOv3 features with local pooling at instrument landmarks from fine-tuned SAM 3.1 masks. Our shared Transformer encoder and temporal convolutional heads integrate this representation with mask geometry, monocular depth and visual state estimates from arm-specific multilayer regression networks. Our position branch predicts displacement magnitude and direction separately to preserve traveled distance. We fit trajectories to predicted state observations and motion increments by differentiable weighted least squares, expressing quaternion observations relative to cumulative predicted rotations to obtain a quadratic orientation objective. We evaluate reconstruction across 2,802 Open-H episodes. Compared with LiveMAE on the main Open-H benchmark, our method reduces path-length mean absolute error from 0.45 to 0.34,cm and increases temporal mean average precision for motion segmentation from 44.54% to 54.44%.
Purpose: Marker-based tracking of surgical robots is occlusion-prone in cluttered operating rooms. We evaluate stereo differentiable rendering for marker-free, real-time robot pose tracking, potentially improving safety, reducing setup time, and enabling multi-robot interaction. Methods: We extend the markerless pose estimation framework roboreg to online dynamic tracking via (i) sequential optimisation that propagates pose estimates across frames with motion-adaptive hyperparameter tuning, and (ii) CUDA stream parallelisation of segmentation and optimisation, combined with CUDA-graph accelerated segmentation. We evaluate on 38 unobstructed and 5 occluded displacement sequences with static start/end ground-truth calibrations and dynamic marker-based reference tracking. Results: We achieve real-time 1080p tracking at 30 fps (up from 14 fps for vanilla roboreg), matching the camera frame rate. Accuracy reaches 1.7 cm / 0.6 deg against static ground truth and 1.2 cm mean 3D error over 27,460 frames against the marker-based reference (1.53 cm over 1,242 occluded frames). Our method outperforms FoundationPose by 11% in dynamic estimation (63% under occlusion) and 250% in static estimation, with 6x faster inference. Conclusions: Stereo differentiable rendering enables real-time, high-resolution marker-free surgical robot tracking, on par with marker-based approaches and surpassing foundation-model baselines.
Yanghe Hao, Martin Huber, Christos Bergeles +1
School of Biomedical Engineering & Imaging Sciences, King’s College London, London, United Kingdom.