Onboard Vision and MPC Navigation for Underwater Robots: An Open BlueROV2 Platform for Multi-Robot Experiments & Docking
Authors: Victor Nan Fernandez-Ayala, Wiktor Kowalczyk, Cezary Banaszek, Dimos V. Dimarogonas
Organizations: Department of Decision and Control Systems, School of Electrical Engineering and Computer Science, KTH Royal Institute of Technology, Stockholm, Sweden.
Autonomous underwater robots require robust perception, estimation and control to operate in confined environments. This paper presents an open-source BlueROV2 platform combining onboard vision with nonlinear Model Predictive Control (NMPC) for autonomous navigation and docking. The platform integrates an NVIDIA Jetson Orin NX and an Intel RealSense D435i stereo camera in a modular pressure housing. Underwater-calibrated stereo depth and realtime object detection provide relative position measurements of nearby BlueROV2 vehicles in the camera and body frames. A quaternion-based estimator fuses external pose and inertial measurements, while an NMPC controller based on a nonlinear six-degree-of-freedom model tracks planned navigation and docking trajectories. To support reproducible development, we also provide open-source physics-based PX4 SITL and Gazebo environments, multi-robot simulation tools and a lowcost physical docking station. Experiments evaluate underwater perception, onboard computational performance, state estimation, trajectory tracking and autonomous docking.
Figures & tables
Fig. 1 : BlueROV2 Heavy with the modular autonomy tube.
TABLE I : Parameters of the modified BlueROV2 Heavy.
Fig. 2 : Underwater BlueROV2 detection and aligned stereo-depth output produced by the onboard perception module.
Fig. 3 : Open experimental infrastructure. The physical docking station and the simulator from the multi-robot launch file.
Parameter
Value
Camera profiles
RGB 1280x720x15fps; D 848x480x30fps
YOLO input resolution
1280x720 pixels
YOLO confidence threshold
0.9
Jetson power mode
15W
IMU sampling rate
200 Hz
Motion-capture rate
100 Hz
TABLE II : Experimental and algorithmic parameters.
Fig. 4 : Perception evaluation: (a) stereo range versus motion-capture ground truth; (b) held-out YOLO precision, recall and mAP; and (c) Jetson Orin NX utilization and inference time for the selected profile.
Fig. 5 : Docking evaluation: (a) raw motion-capture and EKF surge velocity during the final hold; (b) measured horizontal trajectory, commanded dock pose and final position; and (c) overlaid frames of the physical docking sequence.
Vision-based underwater target tracking is challenged by unreliable depth measurements and unknown target motion. This paper proposes a stereo visual-servoing framework for an autonomous underwater vehicle (AUV). For perception, the framework derives a stable 3D relative state from stereo images through target-specific depth extraction and Kalman filtering. It constructs a target-depth mask from color, disparity, and temporal cues to select reliable target pixels, and then filters the resulting depth measurement and detected image center separately. For control, the framework decouples yaw regulation from translational control, avoiding computationally expensive coupled multi-DOF optimization and enabling real-time translational MPC. The translational controller employs adaptive model-fusion predictive control, combining constant-velocity and zero-velocity target models to accommodate different target-motion patterns. It updates the model weights using historical prediction errors and computes translational commands subject to actuation, following-distance, and field-of-view constraints. Through simulations and real-world experiments, we validate the effectiveness of the proposed framework and show it has better performance than existing frameworks.
Yuheng Zhou, Haiyang Cheng, Yanqi Feng +5
Department of Automation, Shanghai Jiao Tong University, Shanghai, China · University of Toronto, Toronto, Canada · University of Calgary, Calgary, Canada
This paper presents a fully integrated vision-based framework for real-time and robust localization, autonomous navigation, and mapping for unmanned underwater vehicles (UUVs) in dynamic, visually challenging environments. The proposed pipeline enables both net-relative and global localization while generating continuous 3D maps of the surroundings in real-time. The framework was validated on synthetic datasets with ground truth and tested onboard an UUV during autonomous net-relative navigation experiments. Results demonstrate real-time performance and enhanced robustness, supporting vision-driven autonomous navigation and enabling the field deployment of marine robots for critical inspection and mapping tasks in complex underwater environments.
Erik Tjærand Frøland, Marco Job, Md Ether Deowan +1
Dept. of Mechanical and Industrial Engineering, NTNU, Norway · Oceaneering AS, Norway
Underwater robots rely on complementary sensors whose reliability changes abruptly with water visibility and vehicle motion. We introduce AquaJEPA, a sensor-configurable family of action-conditioned joint-embedding predictive models spanning full multimodal, camera-only, sonar-only, and sensor-dropout configurations. Its members share a latent objective and receding-horizon control interface that predict future representations and physical dynamics from camera, forward-looking sonar, proprioception, and thruster commands. Trained from scratch on one hour of action-labelled data, the family is evaluated in Stonefish on 120 fresh paired scenarios spanning unseen layouts, visibility changes, dynamics shifts, and scheduled DVL loss. AquaJEPA-base achieves the strongest aggregate closed-loop performance, improving success over state-only by 12.5 percentage points and reducing final error by 0.189 m; both paired 95% intervals exclude zero. In a separate three-seed evaluation, it reduces paired final error relative to AquaJEPA-S by 0.118 m, with the same direction for every seed. AquaJEPA-robust more than halves prediction error during camera and camera-DVL blackouts. These results show that full multimodal prediction improves over state-only control and the sonar-only family member in this benchmark, while sensor-dropout training provides robustness under sensor loss.