Recon2Servo: Robotic Ultrasound Visual Servoing via Learned Image-to-Motion Inference
Authors: Yameng Zhang, Pei Liu, Dianye Huang, Yizhao Qian, Zhongyu Chen, Xiangyu Chu, K. W. Samuel Au, Zhongliang Jiang
Organizations: Department of Mechanical Engineering, The University of Hong Kong, Hong Kong SAR. · Department of Mechanical and Automation Engineering, The Chinese University of Hong Kong, Hong Kong SAR. · School of Computation, Information and Technology, Technical University of Munich, Germany. · Department of Electronic Engineering, The Chinese University of Hong Kong, Hong Kong SAR. · School of Optometry, Hong Kong Polytechnic University, Hong Kong SAR.
Ultrasound visual servoing is essential for autonomous robotic ultrasound, yet 6-DoF probe control from 2D B-mode images remains challenging due to limited and ambiguous out-of-plane motion cues. Existing methods typically rely on anatomical priors or handcrafted visual features, limiting their generalizability across imaging targets. Inspired by trackerless 3D ultrasound reconstruction, we propose Recon2Servo, a visual servoing framework that learns image-to-motion inference directly from B-mode images for 6-DoF probe control. A DINOv3 encoder with low-rank adaptation and a bidirectional relation module estimate the relative probe pose between current and target images to guide iterative closed-loop target-view alignment. The framework combines supervised relative-pose learning, reconstruction-guided closed-loop adaptation, and bounded residual pose correction to improve motion inference during servoing. Evaluations on a public dataset and an in-house dataset collected from 12 healthy volunteers using different ultrasound systems demonstrate its effectiveness in reconstructed-volume servoing. Additional real-robot demonstrations of target-view alignment and dynamic tracking on a human forearm are provided in the supplementary video: https://youtu.be/qwsOdI-GMYk.
Figures & tables
Fig. 1: Overview of the proposed ultrasound servoing framework. Tracked B-mode images are reconstructed into a 3-D volume, which serves as an interactive simulation environment for training and adapting the probe control policy. The resulting policy is deployed using real-time ultrasound image feedback to iteratively guide the probe toward a target view.
Fig. 2: Relative-pose estimation network. A shared LoRA-adapted DINOv3 encoder and bidirectional relation module map current and target images to translation and rotation predictions. Snowflakes and flames denote frozen and trainable components, respectively.
Fig. 3: Reconstruction-guided closed-loop adaptation. Tracked images and calibrated poses define the reconstruction environment. A frozen pose estimator drives fixed-gain rollouts, and the renderer returns images at the updated poses. The known current and target poses provide oracle labels for visited states, which are combined with recorded training pairs to update the pose estimator.
Fig. 4: Bounded residual pose correction. The adapted encoder, relation module, and pose heads are frozen. The two motion tokens and normalized mean estimate are concatenated and passed to a trainable residual MLP. Its output is bounded in physical twist coordinates and added to the mean.
Fig. 5: Robotic ultrasound acquisition setup for the in-house dataset. A SonoScape L741 linear transducer connected to a SonoScape E2 system is mounted on a Realman RM75 robotic arm for automated forearm scanning.
Fig. 6: Reconstructed ultrasound volumes and image regions used for servoing in (a) the public TUS-REC2024 dataset and (b) the in-house dataset. Left: reconstructed volumes with the imaging planes outlined in red. Right: corresponding B-mode views, with red boxes indicating the 30×30 mm ROIs resized to 256×256 pixels for model input.
Fig. 7: Dataset-specific training protocols. The public-dataset estimator is trained on recorded pairs, adapted using 80% recorded and 20% rollout pairs, and refined using rollout pairs alone. For the in-house dataset, the public-dataset Stage-2 estimator initializes adaptation using labeled in-house rollout pairs, followed by residual-head training on the same rollout data.
Dataset
Model
Overall ↑
2-4 mm ↑
4-7 mm ↑
7-10 mm ↑
10-15 mm ↑
Public
A: Initial estimator
512/720 (71.11%)
86.67%
71.67%
65.00%
61.11%
B: A + DAgger
589/720 (81.81%)
92.22%
80.00%
78.89%
76.11%
C: A + Residual
594/720 (82.50%)
90.56%
81.67%
82.22%
75.56%
D: A + DAgger + Residual
640/720 (88.89%)
95.56%
88.89%
88.89%
82.22%
In-house
Public A (zero-shot)
189/576 (32.81%)
52.1%
38.2%
27.1%
13.9%
Public B (zero-shot)
228/576 (39.58%)
59.03%
44.44%
34.03%
20.83%
TABLE I: Overall and distance-stratified reaching success.
Dataset
Model
Out of range ↓
Step limit ↓
Out of volume ↓
Public
A: Initial estimator
16 (2.22%)
192 (26.67%)
0 (0.00%)
B: A + DAgger
14 (1.94%)
116 (16.11%)
1 (0.14%)
C: A + Residual
65 (9.03%)
61 (8.47%)
0 (0.00%)
D: A + DAgger + Residual
30 (4.17%)
50 (6.94%)
0 (0.00%)
In-house
Public A (zero-shot)
170 (29.51%)
199 (34.55%)
18 (3.13%)
Public B (zero-shot)
154 (26.74%)
177 (30.73%)
17 (2.95%)
TABLE II: Failure counts and percentages by termination category.
Freehand 3-D ultrasound (US) imaging has attracted increasing attention owing to its intuitive volumetric visualization, ease of use, and low cost. However, accurate 3-D reconstruction critically depends on stable probe pose estimation, yet existing trackerless methods remain susceptible to accumulated pose errors, particularly over long scanning trajectories. To address this limitation, we propose a global-to-local pose estimation framework that exploits external camera observations for globally stable localization and B-mode US images for anatomy-aware local refinement. Specifically, the framework comprises a dual-camera branch that performs contextual feature aggregation across camera views and temporal observations to estimate a globally consistent probe trajectory, and a B-mode branch that performs anatomical feature aggregation from sequential US images to capture tissue-dependent local motion cues. A cross-modal fusion module subsequently integrates the contextual camera features and anatomical US features to predict pose residuals and refine the camera-derived estimates in the transformation space. Furthermore, a multi-scale pose loss constrains relative motion over multiple temporal horizons to suppress accumulated drift during extended scans. The proposed framework is validated on phantom and in vivo datasets. On two in-house datasets (FUSION-J and FUSION-L) collected using different machines, the proposed US + Dual-Cam model reduces average trajectory drift to 1.67 mm and 1.29 mm, representing improvement of 16.50% and 27.12%, respectively, over a strong dual-camera baseline, while substantially outperforming US-only pose estimation (>13 mm drift). In in vivo forearm arteries reconstruction, it achieves Hausdorff distances of 1.58 mm, demonstrating the effectiveness of the proposed method on real clinical scenarios.
Yameng Zhang, Zhongyu Chen, Dianye Huang +3
Department of Mechanical Engineering, The University of Hong Kong, Hong Kong SAR, China · Department of Mechanical and Automation Engineering, The Chinese University of Hong Kong, Hong Kong SAR, China
Safe robot-assisted ultrasound imaging requires a reliable controller able to detect and localize probe--tissue interaction. In this paper, we present a B-mode ultrasound image-based contact perception method and a contact-aware impedance controller for robotic ultrasound imaging. The proposed method detects acoustic contact independently of force measurements, enabling contact-conditioned force/torque taring to reduce residual wrench bias. During contact, the method continuously estimates the effective contact location along the curved probe surface and uses it to update the controller interaction frame, enabling visual servoing of the physical probe--tissue contact point during imaging. Experiments on an agar phantom demonstrated a contact-localization RMSE of 1.46±0.14~mm over probe roll angles from −15∘ to 15∘. During static rolling, the proposed controller maintained task-space tracking accuracy comparable to a conventional fixed-frame impedance controller while reducing the maximum compressive interaction force from 31.56~N to 20.09~N, corresponding to a 36.3% reduction. These results demonstrate the potential of ultrasound images as direct contact feedback for safe and accurate robot-assisted ultrasound imaging.
MD Miraj Arefin, M Efe Tiryaki
Center for Robotics and AI (ROMER), Middle East Technical University, Ankara, Turkey
Ultrasound (US) provides real-time, radiation-free imaging, but the image quality depends strongly on how the probe is oriented against the patient body. Robotic US can reduce operator workload and improve acquisition consistency; however, most existing systems focus on normal positioning, where the probe is maintained perpendicular to the local surface. This constraint is inadequate for examinations like echocardiography, where obtaining a diagnostic view requires a non-normal probe angle. Consequently, a clinically useful robotic system must sense the local surface in real-time and preserve the desired probe orientation. Here, we propose an omni-directional probe-orientation control framework that integrates RGB-D perception, local-surface modeling, and task-space orientation control. The surface model fuses multi-view point clouds and provides a quadratic estimate of the local surface. A desired imaging direction is then encoded relative to the normal, enabling the probe to track arbitrary angles. The framework was evaluated through flat-surface tracking, phantom target-angle recovery, and in-vivo tracking of an expert selected view. Results show that the mean angular tracking error was 1.06 +- 0.66 deg. The system recovered a non-normal tilt angle of up to 44.39 +- 2.59 deg relative to the surface normal, and acquired the desired heart chamber view in the phantom and in-vivo experiments.
Xihan Ma, Haichong Zhang
Department of Robotics Engineering, Worcester Polytechnic Institute, 100 Institute Road, Worcester, MA 01609 USA