Autonomous Surface Vehicles (ASVs) operating in dynamic marine environments require robust control policies for tasks such as path following and station keeping, making reinforcement learning (RL) a promising alternative to classical controllers. However, existing ASV simulators rarely support parallel environments for RL training. Such existing simulators require accurate hydrodynamic modeling from computational fluid dynamics solvers or towing tank tests for setting hydrodynamic parameters to address the sim-to-real gap. To address these challenges, we present an ASV simulator and accompanying pipeline that enables training policies starting from unknown vehicle dynamics. Our framework uses only a CAD model and brief set of open-water field trajectories for approximating and refining both hydrodynamic and thruster parameters. Real-world deployments on a BlueBoat ASV demonstrate successful zero-shot sim-to-real transfer in path following and station-keeping tasks without prior hydrodynamic and propeller information.
Figures & tables
Fig. 1: Overview of our sim-to-real pipeline. Initial hydrodynamic parameters are derived from a hull CAD model via a boundary element method (BEM) solver alongside a set of open-water trajectories from the physical autonomous surface vehicle (ASV). System identification (SysID) is performed in the simulator using the covariance matrix adaptation evolution strategy (CMA-ES) to obtain refined hydrodynamic ( Xu,Xuu,Yv,Yvv,Nr,Nrr,Iz ) and thruster ( cu,crev ) parameters, grounding the simulator to the real ASV. Reinforcement learning (RL) policies are trained in the grounded simulator, enabling zero-shot transfer for path following and station-keeping tasks.
Parameter
Port
Starboard
Units
Reverse Thrust Limit ( fmin,i )
-10.0
-10.0
N
Forward Thrust Limit ( fmax,i )
20.0
20.0
N
Motor Force Deadband ( fdead )
5.0
N
Motor Spool-up Lag ( τlag )
0.15
s
Thrust advance loss ( cu )
0.6016
—
Reverse thrust coefficient ( crev )
0.4356
—
TABLE I: Nominal Action Space Parameter Values
Reward Component
wk
Hyperparameters
Cross-Track Error ( rtcross )
2.0
σcross=0.3
PF Heading ( rtPF,head )
1.0
σPF,head=0.3
Path Progress ( rtprog )
1.5
σprog=0.25,utarget=0.5
PF Term Penalty ( rtPF,term )
−10.0
dtermPF=2.0
Target Distance ( rtdist )
2.0
σdist=0.3
SK Heading ( rtSK,head )
1.0
σSK,head=0.3,ϵSK,head=2∘
TABLE II: Reward Function Weights and Hyperparameters
Parameter
Distribution
Hydrodynamic multiplier
U(0.9,1.1)
Advance loss ( cu ) multiplier
U(0.9,1.1)
Reverse efficiency ( cu ) multiplier
U(0.9,1.1)
Motor spool-up ( τlag )
U(0.10,0.20)
Thrust asymmetry ( cstbd )
U(0.95,1.05)
PF wind speed (m/s)
Curriculum: U(0,0)→U(0,5)
TABLE III: Domain Randomization Parameters
Parameter
Path Following (PF)
Station Keeping (SK)
Episode Duration
20 s
30 s
Rollout Steps
128
256
Discount Factor ( γ )
0.99
0.9975
TABLE IV: Task-Specific Training Hyperparameters
Fig. 2: Real-world path following trajectory evaluations for each configuration: Real-World (top-left), Proposed (top-right), Zigzag (bottom-left), and CMA-ES DR (bottom-right). Proposed achieves the lowest average cross-track RMSE and single-trajectory grounding suffers from divergent dynamics.
Parameter
Real-World
Proposed
Zigzag
Units
Iz
4.86
6.31
4.96
kg ⋅ m 2
Xu
-0.46
-0.59
-0.59
kg/s
Xuu
-6.83
-6.98
-8.87
kg/m
Yv
-54.89
-71.26
-71.28
kg/s
Yvv
-114.30
-148.37
-148.42
kg/m
Nr
-6.38
-4.48
-8.29
kg ⋅ m 2 /s
TABLE V: SysID Parameters Across Experimental Configurations
Fig. 3: Real-world station-keeping evaluations lasting 45 s for each configuration: Real-World (top-left), Proposed (top-right), Zigzag (bottom-left), and CMA-ES DR (bottom-right). Grounded policies all similarly outperform Real-World, with Zigzag notably outperforming CMA-ES DR.