Autonomous Surface Vehicles (ASVs) operating in dynamic marine environments require robust control policies for tasks such as path following and station keeping, making reinforcement learning (RL) a promising alternative to classical controllers. However, existing ASV simulators rarely support parallel environments for RL training. Such existing simulators require accurate hydrodynamic modeling from computational fluid dynamics solvers or towing tank tests for setting hydrodynamic parameters to address the sim-to-real gap. To address these challenges, we present an ASV simulator and accompanying pipeline that enables training policies starting from unknown vehicle dynamics. Our framework uses only a CAD model and brief set of open-water field trajectories for approximating and refining both hydrodynamic and thruster parameters. Real-world deployments on a BlueBoat ASV demonstrate successful zero-shot sim-to-real transfer in path following and station-keeping tasks without prior hydrodynamic and propeller information.
Figures & tables
Fig. 1: Overview of our sim-to-real pipeline. Initial hydrodynamic parameters are derived from a hull CAD model via a boundary element method (BEM) solver alongside a set of open-water trajectories from the physical autonomous surface vehicle (ASV). System identification (SysID) is performed in the simulator using the covariance matrix adaptation evolution strategy (CMA-ES) to obtain refined hydrodynamic ( Xu,Xuu,Yv,Yvv,Nr,Nrr,Iz ) and thruster ( cu,crev ) parameters, grounding the simulator to the real ASV. Reinforcement learning (RL) policies are trained in the grounded simulator, enabling zero-shot transfer for path following and station-keeping tasks.
Parameter
Port
Starboard
Units
Reverse Thrust Limit ( fmin,i )
-10.0
-10.0
N
Forward Thrust Limit ( fmax,i )
20.0
20.0
N
Motor Force Deadband ( fdead )
5.0
N
Motor Spool-up Lag ( τlag )
0.15
s
Thrust advance loss ( cu )
0.6016
—
Reverse thrust coefficient ( crev )
0.4356
—
TABLE I: Nominal Action Space Parameter Values
Reward Component
wk
Hyperparameters
Cross-Track Error ( rtcross )
2.0
σcross=0.3
PF Heading ( rtPF,head )
1.0
σPF,head=0.3
Path Progress ( rtprog )
1.5
σprog=0.25,utarget=0.5
PF Term Penalty ( rtPF,term )
−10.0
dtermPF=2.0
Target Distance ( rtdist )
2.0
σdist=0.3
SK Heading ( rtSK,head )
1.0
σSK,head=0.3,ϵSK,head=2∘
TABLE II: Reward Function Weights and Hyperparameters
Parameter
Distribution
Hydrodynamic multiplier
U(0.9,1.1)
Advance loss ( cu ) multiplier
U(0.9,1.1)
Reverse efficiency ( cu ) multiplier
U(0.9,1.1)
Motor spool-up ( τlag )
U(0.10,0.20)
Thrust asymmetry ( cstbd )
U(0.95,1.05)
PF wind speed (m/s)
Curriculum: U(0,0)→U(0,5)
TABLE III: Domain Randomization Parameters
Parameter
Path Following (PF)
Station Keeping (SK)
Episode Duration
20 s
30 s
Rollout Steps
128
256
Discount Factor ( γ )
0.99
0.9975
TABLE IV: Task-Specific Training Hyperparameters
Fig. 2: Real-world path following trajectory evaluations for each configuration: Real-World (top-left), Proposed (top-right), Zigzag (bottom-left), and CMA-ES DR (bottom-right). Proposed achieves the lowest average cross-track RMSE and single-trajectory grounding suffers from divergent dynamics.
Parameter
Real-World
Proposed
Zigzag
Units
Iz
4.86
6.31
4.96
kg ⋅ m 2
Xu
-0.46
-0.59
-0.59
kg/s
Xuu
-6.83
-6.98
-8.87
kg/m
Yv
-54.89
-71.26
-71.28
kg/s
Yvv
-114.30
-148.37
-148.42
kg/m
Nr
-6.38
-4.48
-8.29
kg ⋅ m 2 /s
TABLE V: SysID Parameters Across Experimental Configurations
Fig. 3: Real-world station-keeping evaluations lasting 45 s for each configuration: Real-World (top-left), Proposed (top-right), Zigzag (bottom-left), and CMA-ES DR (bottom-right). Grounded policies all similarly outperform Real-World, with Zigzag notably outperforming CMA-ES DR.
Autonomous surface vessels for floating-waste removal operate under varying hydrodynamics, external disturbances, and challenging water-surface perception. We present a field-validated system that combines camera-based polarimetric perception with a lightweight DRL-based controller for floating-waste detection and capture. Camera detections are converted into water-surface target points and tracked by a controller trained entirely in simulation and deployed directly on a retrofitted ASV platform. Our main contribution is a sim-to-real testing methodology that combines a two-stage simulation protocol with a perception abstraction module designed to mimic real camera behavior, enabling reproducible field trials and explicit evaluation of the sim-to-real gap. We apply this framework in matched simulation and field experiments across 14 disturbance regimes to expose failure modes and evaluate robustness. The results show centimeter-level terminal accuracy and indicate robust control performance under the evaluated perturbation regimes. The main source of degradation is insufficient actuation-model fidelity. We also demonstrate the system in a search-and-capture application using real camera detections in real-world conditions over areas of up to 450m2. The study distills practical lessons for reliable transfer, including improved actuation-model fidelity, targeted domain randomization, and careful management of latency and timestamps across modules, while highlighting remaining challenges.
Luis F. W. Batista, Stéphanie Aravecchia, Cédric Pradalier
GeorgiaTech Europe - IRL2958 GT-CNRS, Metz, France
Holonomic autonomous underwater vehicles (AUVs) have the hardware ability for agile maneuvering in both translational and rotational degrees of freedom (DOFs). However, due to challenges inherent to underwater vehicles, such as complex hydrostatics and hydrodynamics, parametric uncertainties, and frequent changes in dynamics due to payload changes, control is challenging. Performance typically relies on carefully tuned controllers targeting unique platform configurations, and a need for re-tuning for deployment under varying payloads and hydrodynamic conditions. As a consequence, agile maneuvering with simultaneous tracking of time-varying references in both translational and rotational DOFs is rarely utilized in practice. To the best of our knowledge, this paper presents the first general zero-shot sim2real deep reinforcement learning-based (DRL) velocity controller enabling path following and agile 6DOF maneuvering with a training duration of just 3 minutes. Sim2Swim, the proposed approach, inspired by state-of-the-art DRL-based position control, leverages domain randomization and massively parallelized training to converge to field-deployable control policies for AUVs of variable characteristics without post-processing or tuning. Sim2Swim is extensively validated in pool trials for a variety of configurations, showcasing robust control for highly agile motions.
Lauritz Rismark Fosso, Herman Biørn Amundsen, Marios Xanthidis +1
Autonomous surface vehicles vary widely in hydrodynamic and actuation characteristics, yet most controllers are designed for single-platform deployment. We present an adaptive reinforcement learning approach for trajectory tracking that enables zero-shot cross-platform deployment using a single policy. Since the deployment platform's dynamics are unknown to the policy, we address cross-platform generalization with the standard partial-observability approach of conditioning on interaction history, employing a teacher-student architecture in which a learned module infers a latent representation of the platform dynamics. The policy is trained in simulation under randomized vessel dynamics and is deployed zero-shot to two real-world platforms without any fine-tuning, despite relying on a simple analytical dynamics model rather than a high-fidelity hydrodynamic simulator. In real-world experiments on two different platforms, the adaptive policy outperforms non-adaptive learning-based baselines by up to 58% in position mean absolute error while approaching the tracking accuracy of a platform-specific tuned controller.
Ruiheng Jiang, Thomas Bi, Raffaello D'Andrea +1
Institute for Dynamic Systems and Control, ETH Zurich, Switzerland.