Fading Expert Guidance: Bridging Model-Based and Learning-Based Control for Abortable Autonomous Overtaking
Organizations: Intelligent Robotics Group, Department of Electrical Engineering and Automation, Aalto University, 02150 Espoo, Finland
Abstract
Overtaking on two-lane roads is a safety-critical decision-making problem for autonomous vehicles, since oncoming traffic may force the ego vehicle to abort and merge back to its original lane. Deep reinforcement learning (DRL) is promising for such continuous-control tasks, but it requires substantial interaction data and may produce unsafe exploratory actions during training. Model-based controllers provide structured behavior, but depend on model fidelity and hand-designed logic. This paper proposes a fading expert-guidance framework that uses a model-based controller to guide DRL training for autonomous overtaking. The expert combines a constrained iterative LQR (CiLQR) planner with PID-based auxiliary controllers, activated when the optimized plan becomes infeasible. The expert action enters the actor objective through a fading term, interpreted as a time-varying soft trust region around the expert policy. This term first biases the policy toward the expert and then vanishes, allowing optimization according to the reinforcement-learning objective. The method is evaluated with PPO, TD3, and SAC. Simulations with bidirectional traffic show improved sample efficiency and final task performance. The safety effects are algorithm-dependent: guidance reduces vehicle collisions for SAC, eliminates boundary collisions for TD3, and leaves the PPO collision profile unchanged. Overall, the results support fading expert guidance as a practical mechanism for transferring model-based driving knowledge to learning-based controllers without restricting the final policy to imitation in an abortable overtaking task with bidirectional traffic.
Figures & tables
| Driver types (m/s 2 ) | Desired speed (m/s) | Desired time (s) | Jam distance (m) | Max acceleration (m/s 2 ) | Desired deceleration (m/s 2 ) | Politeness acceleration (m/s 2 ) | Safe braking (m/s 2 ) | Acceleration threshold (m/s 2 ) | Length acceleration (m) | Width acceleration (m) |
| Timid | 27.8 | 2.0 | 4.0 | 0.8 | -1.0 | 1.0 | 1.0 | 0.2 | 5 | 2 |
| Normal | 33.3 | 1.5 | 2.0 | 1.4 | -2.0 | 0.5 | 2.0 | 0.1 | 5 | 2 |
| Aggressive | 38.9 | 1.0 | 0.0 | 2.0 | -3.0 | 0.0 | 3.0 | 0.0 | 5 | 2 |
| Truck | 23.6 | 2.0 | 4.0 | 0.7 | -2.0 | 1.0 | 1.0 | 0.2 | 6 | 2.5 |
| Evaluation Metrics | Guidance System | PPO | G-PPO | TD3 | G-TD3 | SAC | G-SAC |
|---|---|---|---|---|---|---|---|
| Episode reward | 5940.9 | 5437.31 | 6421.85 | 3462.42 | 6702.52 | 10112.66 | 14879.58 |
| Speed (m/s) | 42.97 | 40.1 | 44.76 | 45.85 | 45.19 | 59.64 | 57.2 |
| Displacement (m) | 440.26 | 350.29 | 466.68 | 356.76 | 451.07 | 638.42 | 807.73 |
| Computation time (ms) | 74 | 0.1 | 0.33 | 0.82 | 0.27 | 0.89 | 0.86 |
| Energy consumption | 1.21 | 0.34 | 1.39 | 2.45 | 0.98 | 1.36 | 0.52 |
| Vehicles collision rate | 1 | 1 | 1 | 0.2 | 0.8 | 0.7 | 0.3 |