Efficient and robust tracking of surgical robotic instruments is important for robot-assisted minimally invasive surgery, yet remains challenging due to the complexity of surgical scenes and the unconventional geometry of surgical instruments. Keypoint-based approaches are efficient, but their performance depends on reliable feature detection. Improving these detectors with real-world supervision is difficult because accurate real-world annotations are costly to obtain at scale. To address this limitation, we introduce a tracker-guided self-training framework that adapts a model pretrained on synthetic images to unlabeled real-world videos. Given measured robot joint states, an uncertainty-aware EKF recursively corrects the instrument pose and the observable joint angles by comparing projected model features with detected keypoints, shaft boundaries, and mask-derived cues. An RTS smoother subsequently refines the resulting trajectory, which is projected into pseudo-labels for fine-tuning the feature detector without laborious pose annotations. Experiments on real-world videos demonstrate consistent improvements from self-training across all evaluated keypoint metrics, and the resulting model outperforms prior approaches in both accuracy and runtime. The code and data will be released upon publication.
Figures & tables
Fig. 1: Self-training improves sim-to-real feature detection. Pretrained and self-trained models are compared on unseen real-world images with diverse orientations, distractions, and occlusions. Circles and dashed white lines denote manually labeled ground truth; plus signs and solid yellow lines denote model predictions. A tracker-guided self-training framework enables a synthetically pretrained model to bridge the gap to challenging, real-world scenes without tedious human labeling.
Fig. 2: Overview of the proposed self-training framework. A query-based multitask model predicts keypoint heatmaps, shaft edge heatmaps, and instrument masks from RGB video, which are decoded into geometric observations with associated uncertainty. Sequential multi-cue fusion combines these observations with robot kinematics to obtain smoothed pose trajectories, using an initial base-to-camera transform Tbc estimated by differentiable rendering [ 8 ] . An EKF propagates the instrument state between frames and updates it using the current frame’s observations, while a backward RTS smoother refines these pose estimates to produce a smoothed pose trajectory. Pseudo-label generation projects and renders the smoothed instrument configurations into feature labels for sim-to-real adaptation.
Fig. 3: Synthetic data generation. Isaac Sim generates labeled images with ground-truth keypoints and shaft edge annotations.
Metric
Pretrained
Ours
Improvement
Per-keypoint accuracy (PCK@20)
Gripper 1
91.59
93.70
+2.11 pp
Gripper 2
94.92
95.99
+1.07 pp
End effector
94.13
98.12
+3.99 pp
Pitch
82.76
91.94
+9.18 pp
Pitch side
94.60
99.90
+5.30 pp
TABLE I: Keypoint Detection Accuracy. Pretrained and self-trained detectors are evaluated on 16 SurgPose trajectories, with errors measured at the original 1400 × 986 resolution and PCK reported as percentages. The best results are shown in bold.
Model
Gripper (px) ↓
All Kpts. (px) ↓
Edge ↑
Unified detector [ 7 ]
22.24
–
0.9606
Pretrained
17.19
12.57
0.9886
Ours
14.90
11.38
0.9887
TABLE II: Real-World Feature Detection Accuracy. Pretrained and self-trained detectors are evaluated on the annotated real-world data, with mean keypoint errors reported in pixels and edge detection measured by the modified EA score [ 7 ] . The best and second-best results are shown in bold and underlined, respectively.
Observations
Mean error ↓ (px)
PCK@20 ↑ (%)
Kpts.
Shaft
Mask
EKF
RTS
EKF
RTS
✓
–
–
11.50
10.68
91.63
93.99
✓
✓
–
11.08
10.34
92.91
95.78
✓
✓
✓
10.05
9.22
94.36
96.58
TABLE III: Ablation of Multi-Cue Fusion. Shaft and mask cues are incrementally incorporated into pseudo-label generation using the pretrained detector on the training split. Mean keypoint error and PCK@20 are reported for both EKF and RTS. The best and second-best results are shown in bold and underlined, respectively.
Fig. 4: Qualitative tracking comparison. Representative results across different ex vivo backgrounds for a rendering-based tracker [ 10 ] and the proposed EKF tracker with observations from the pretrained and self-trained detectors. The estimated poses are rendered as green silhouettes, while the reference poses are highlighted by white skeletons. The self-trained detector provides more accurate alignment across challenging tissue appearances and unseen backgrounds compared to the rendering-based baseline and the pretrained model.
Method
Mean err. ↓
P95 err. ↓
PCK@20 ↑
FPS ↑
CMA-ES + KF [ 10 ]
20.53
60.47
69.17
28.78
Pretrained (w/ mask)
9.40
22.15
93.13
81.14
Pretrained (w/o mask)
9.77
22.32
92.96
178.64
Ours (w/ mask)
9.27
20.60
94.44
81.08
Ours (w/o mask)
9.23
20.33
94.72
179.07
TABLE IV: Tracking Accuracy and Speed. Results are reported over the 16 SurgPose evaluation trajectories. The best and second-best results are shown in bold and underlined, respectively.