Energy-Aware Path Following: Comparative Analysis of Reinforcement Learning and NMPC for Electric Vehicles
Organizations: Computer Science Department, Najran University, Najran, Saudi Arabia · Institute of Applied Mathematics and Scientific Computing, University of the Bundeswehr Munich, Neubiberg, 85579, Germany
Abstract
Path-following control strategies typically follow the bi-objective optimization dilemma: minimizing deviations from a reference path while maintaining smooth speed profiles. The latter objective is especially relevant for Electric Vehicles (EVs), since their limited driving range can be extended by recovering energy through regenerative braking, a feature that has not yet been sufficiently studied in the literature. In this work, we perform a comparative analysis of four controllers under one common Frenet frame-based kinematic vehicle model, utilizing a validated energy model (VT-CPEM) with explicit regenerative braking. Herein, we implement the following controllers: Nonlinear Model Predictive Control (NMPC), Proximal Policy Optimization (PPO), gain-scheduled Ackermann state-feedback baseline (PID-SF), and a Stanley geometric baseline. To satisfy real-time requirements, we implement the NMPC using JIT-compiled CasADi. Moreover, we train the PPO using traditional straight and S-curve tracks, after which we successfully transfer the unmodified policy to unseen tracks, including: an ISO 3888-1 lane-change, a chicane, randomly-generated parameterized-splines, and a graded road. In addition, the policy transfers to a dynamic single-track vehicle model with linear tires, zero-shot with an acceptable initial performance, which was optimized after brief fine-tuning. Thereby, we demonstrate that our PPO is readily transferable to more comprehensive vehicle models. We conclude with a performance analysis of developed controllers and discuss ideas for future work.
Figures & tables
| Research and Scope | Vehicle Model | Control Strategy | Energy Model | Experiments |
| Fu et al. (2022) : Path tracking at handling limits | Nonlinear vehicle dynamics with tire-force, load-transfer, and adhesion effects | NMPC | - | HIL and simulation using CarSim. |
| Domina and Tihanyi (2022) : Automated path following near handling limits | Linear time-varying vehicle model considering steering dynamics | LTV-MPC | - | Simulation-based using MATLAB. |
| Reiter et al. (2023) : Obstacle avoidance and path following | Kinematic vehicle model in combined Cartesian/Frenet coordinate representation | NMPC | - | Simulation-based using ACADOS. |
| Belkebir et al. (2026) Curvature-aware path tracking with embedded NMPC | Augmented kinematic model in Frenet frame | NMPC with curvature-aware speed and Lyapunov terminal cost | - | Simulation using CARLA and embedded timing with CasADi on Raspberry Pi 5 and Xavier AGX. |
| Li et al. (2019) : Vision-based lateral control | TORCS / VTORCS simulator with a perception-to-control pipeline | RL + deep learning perception | - | Simulator-based training and evaluation using VTORCS; compared against LQR, MPC. |
| Hess and Ljungbergh (2021) : Longitudinal and lateral path following | Custom kinematic bicycle model | DDPG | - | Simulation-based using PyTorch. |
| Symbol | Value | Description |
| Vehicle (adopted from VT-CPEM Fiori et al. (2016) ) | ||
| \mathrm{k}\mathrm{g} | Mass | |
| \mathrm{m} | Wheelbase | |
| Drag coefficient | ||
| \mathrm{m} | Frontal area | |
| Rolling resistance | ||
| Phase | Objective | Steps | Description |
| 0 | Longitudinal control | Straight line ( ), speed tracking | |
| 1 | Lateral control | S-curve, lateral introduction \mathrm{m} | |
| 2 | Longitudinal + Lateral | Refining controls of coupled dynamics | |
| 3 | Disturbance recovery | \mathrm{m} applied | |
| 4 | Energy optimization | Refining controls with respect to energy | |
| Total: 450 k steps; five seeds (42–46). | |||
| Variant | Preview observation | Energy reward | \mathrm{m} | Regen (%) |
| base | ||||
| preview | ||||
| energy-reward | ||||
| preview+energy | ||||
| NMPC (reference) | – | – |
| Controller | \mathrm{m} | \mathrm{k}\mathrm{J} | \mathrm{k}\mathrm{J} | \mathrm{k}\mathrm{J} | Solving time \mathrm{ms} | Speedup |
| NMPC (do-mpc) | – | |||||
| NMPC (compiled) | – | – | – | |||
| PPO (preview+e) | ||||||
| PID-SF | ||||||
| Stanley |
| Controller | |||
| NMPC | |||
| PPO | |||
| PID-SF | |||
| Stanley |
| Configuration | \mathrm{m} | \mathrm{k}\mathrm{J} | |||||
| Best tracking | |||||||
| Best energy | |||||||
| Default |
| Controller | \mathrm{m} | \mathrm{k}\mathrm{J} | Speedup | Failures |
| NMPC | ||||
| PPO | ||||
| PID-SF | ||||
| Stanley |
| Controller | S-curve \mathrm{m} | ISO 3888-1 \mathrm{m} | Chicane \mathrm{m} |
| NMPC | |||
| PPO (zero-shot) | |||
| PID-SF | |||
| Stanley |
| Controller | Tracks | \mathrm{m} | Failures |
| NMPC | |||
| PPO | |||
| Stanley | |||
| PID-SF |
| \mathrm{m} | |
| Var. | Description | \mathrm{m} | \mathrm{k}\mathrm{J} |
| A | Full (Ph. 0 1 2 3 4) | ||
| B | Phase 1 only | ||
| C | Phases 1+2 only | ||
| D | Phase 2 only (no curriculum) |
| Controller | (m) | (kJ) | Regen (%) | Status |
| PPO fine-tuned ( k) | ok | |||
| LQR (dynamic model) | ok | |||
| PPO zero-shot | ok | |||
| NMPC (kinematic model) | ok | |||
| PID-SF (kinematic design) | ok | |||
| Stanley (re-tuned) | ok |
Appendix figures & tables1 asset
Supplementary material from the paper’s appendix.
Appendix
| Phase | |||||||
| 0 (longitudinal, straight) | |||||||
| 1 (lateral introduction) | |||||||
| 2 (stopping, full path) | |||||||
| 3 (disturbance recovery) | |||||||
| 4 (energy optimization) |