Energy-Aware Path Following: Comparative Analysis of Reinforcement Learning and NMPC for Electric Vehicles
Authors: Mohamed Sabaa, Mostafa Emam
Organizations: Computer Science Department, Najran University, Najran, Saudi Arabia · Institute of Applied Mathematics and Scientific Computing, University of the Bundeswehr Munich, Neubiberg, 85579, Germany
Path-following control strategies typically follow the bi-objective optimization dilemma: minimizing deviations from a reference path while maintaining smooth speed profiles. The latter objective is especially relevant for Electric Vehicles (EVs), since their limited driving range can be extended by recovering energy through regenerative braking, a feature that has not yet been sufficiently studied in the literature. In this work, we perform a comparative analysis of four controllers under one common Frenet frame-based kinematic vehicle model, utilizing a validated energy model (VT-CPEM) with explicit regenerative braking. Herein, we implement the following controllers: Nonlinear Model Predictive Control (NMPC), Proximal Policy Optimization (PPO), gain-scheduled Ackermann state-feedback baseline (PID-SF), and a Stanley geometric baseline. To satisfy real-time requirements, we implement the NMPC using JIT-compiled CasADi. Moreover, we train the PPO using traditional straight and S-curve tracks, after which we successfully transfer the unmodified policy to unseen tracks, including: an ISO 3888-1 lane-change, a chicane, randomly-generated parameterized-splines, and a ±3∘ graded road. In addition, the policy transfers to a dynamic single-track vehicle model with linear tires, zero-shot with an acceptable initial performance, which was optimized after brief fine-tuning. Thereby, we demonstrate that our PPO is readily transferable to more comprehensive vehicle models. We conclude with a performance analysis of developed controllers and discuss ideas for future work.
Figures & tables
Research and Scope
Vehicle Model
Control Strategy
Energy Model
Experiments
Fu et al. (2022) : Path tracking at handling limits
Nonlinear vehicle dynamics with tire-force, load-transfer, and adhesion effects
NMPC
-
HIL and simulation using CarSim.
Domina and Tihanyi (2022) : Automated path following near handling limits
Linear time-varying vehicle model considering steering dynamics
LTV-MPC
-
Simulation-based using MATLAB.
Reiter et al. (2023) : Obstacle avoidance and path following
Kinematic vehicle model in combined Cartesian/Frenet coordinate representation
NMPC
-
Simulation-based using ACADOS.
Belkebir et al. (2026) Curvature-aware path tracking with embedded NMPC
Augmented kinematic model in Frenet frame
NMPC with curvature-aware speed and Lyapunov terminal cost
-
Simulation using CARLA and embedded timing with CasADi on Raspberry Pi 5 and Xavier AGX.
Li et al. (2019) : Vision-based lateral control
TORCS / VTORCS simulator with a perception-to-control pipeline
RL + deep learning perception
-
Simulator-based training and evaluation using VTORCS; compared against LQR, MPC.
Hess and Ljungbergh (2021) : Longitudinal and lateral path following
Custom kinematic bicycle model
DDPG
-
Simulation-based using PyTorch.
Table 1: Comparative overview of representative control and planning methods discussed in this section.
Figure 1: Reference velocity vref(s) and curvature κ(s) on an S-curve. The terminal deceleration ramp drives regenerative energy recovery.
Symbol
Value
Description
Vehicle (adopted from VT-CPEM Fiori et al. (2016) )
m
1500(\mathrm{k}\mathrm{g})
Mass
Lf
4.0(\mathrm{m})
Wheelbase
CD
0.30
Drag coefficient
Af
2.2(\mathrm{m}2)
Frontal area
Cr
0.01
Rolling resistance
Table 2: Complete Parameter List.
Figure 2: PID-SF unit step response on the kinematic model ( d0=1(\mathrm{m}),v=5(\mathrm{m}\mathrm{s}−1) ).
Figure 3: Closing the regenerative-energy gap. Left: regenerated fraction Er/Et for the four ablation variants against the NMPC reference; annotations highlight the share of the gap closed. Right: the same variants in the tracking–energy plane, signifying that preview improves both objectives simultaneously.
Variant
Preview observation
Energy reward
RMSEd(\mathrm{m})
Regen (%)
base
×
×
0.459
9.3
preview
✓
×
0.310
23.1
energy-reward
×
✓
0.398
32.7
preview+energy
✓
✓
0.314
47.9
NMPC (reference)
–
–
0.474
59.9
Table 4: 2×2 ablation of the two proposed changes (primary seed, S-curve, benchmark IC). Gap closed is measured against NMPC’s regenerative fraction: preview 27 %, energy-reward 46 %, both 76 %. Energies are at the single benchmark IC; N=30 means differ compared to Table 8 .
Figure 4: Lateral deviation d(t) on the S-curve for the four controllers at the benchmark initial condition. Legend values are the RMSE of the plotted tracks.
Controller
RMSEd(\mathrm{m})
Et(\mathrm{k}\mathrm{J})
Er(\mathrm{k}\mathrm{J})
Enet(\mathrm{k}\mathrm{J})
Solving time (\mathrm{ms})
Speedup
NMPC (do-mpc)
0.474
41.69
24.95
16.79
117.6
–
NMPC (compiled)
0.453
–
–
–
16.3
1×
PPO (preview+e)
0.314
49.76
23.85
25.92
0.407
40×
PID-SF
1.155
57.76
24.74
33.02
0.065
251×
Stanley
1.666
57.72
25.24
32.48
0.088
185×
Table 5: Main Results (S-curve, x0=[0,0.5,0.1,0,3]⊤ ). Energies in kJ ; solve time in ms . Speedups are relative to the JIT-compiled NMPC, cf. Section 5.3 . The do-mpc timings are: mean 117.6 , p95 149.5 , max 164.5 . The compiled solver was executed for timing analysis only; its energy figures are omitted because it solves the identical optimal control problem and its trajectory deviation ( 0.453 vs. 0.474 ) differs only by solver tolerance.
Figure 5: Heading error α(t) on the S-curve for the four controllers at the benchmark initial condition. NMPC and PPO preview+energy maintain smooth orientation tracking, while PID-SF and Stanley exhibit considerable oscillations.
Figure 6: Control inputs on the S-curve: steering rate u1=δ˙ (left) and longitudinal acceleration u2=v˙ (right). Both NMPC and PPO strictly respect actuator bounds without high-frequency oscillations.
Figure 7: Kinematic lateral acceleration on the S-curve. The black dashed line marks the dynamics handling limit alat,max .
Figure 8: Energy breakdown per controller. Percentages show regenerative fraction Er/Et . NMPC achieves the highest recovery at 59.9 %, while the preview policy only reaches 47.9 %.
Controller
ηr=0.50
ηr=0.65
ηr=0.80
NMPC
22.55
16.79
11.04
PPO
31.42
25.92
20.42
PID-SF
38.73
33.02
27.31
Stanley
38.30
32.48
26.65
Table 6: Enet(\mathrm{k}\mathrm{J}) versus regeneration efficiency ηr .
Configuration
wd
wα
wv
wu1
wu2
RMSEd(\mathrm{m})
Enet(\mathrm{k}\mathrm{J})
Best tracking
76.50
100.75
6.00
0.81
0.80
0.309
20.59
Best energy
4.33
239.75
0.95
4.95
3.32
0.719
15.84
Default
10.0
100.0
1.0
1.0
1.0
0.474
16.79
Table 7: Extremes of the NMPC weight Pareto front ( N=30 LHS, nine non-dominated configurations).
Figure 9: Pareto-frontier analysis for the NMPC weights with N=30 LHS. The selected (default) configuration approaches the frontier, exhibiting an acceptable outcome without exhaustive tuning.
Controller
RMSEd(\mathrm{m})
Enet(\mathrm{k}\mathrm{J})
Speedup
Failures
NMPC
0.520±0.120
10.34±7.35
1×
0/30
PPO
0.387±0.190
19.36±7.14
250×
0/30
PID-SF
1.337±0.434
27.96±7.55
1880×
5/30
Stanley
1.713±0.127
28.31±8.54
1535×
4/30
Table 8: Statistical results over N=30 LHS initial conditions, which differ from the single-IC values of Table 5 . Wilcoxon signed-rank (paired), PPO < NMPC: p<0.0001 ; Cohen’s d=0.826 . Speedups here are relative to the interpreted do-mpc solver; check Section 5.3 for the compiled comparison.
Table 9: Multi-track evaluation. The policy is not retrained on any evaluation track.
Figure 11: Zero-shot evaluation on the ISO 3888-1 double lane change. The policy was trained only on the S-curve; no retraining was performed.
Controller
Tracks
RMSEd(\mathrm{m})
Failures
NMPC
30
0.163±0.038
0/30
PPO
30
0.256±0.032
0/30
Stanley
30
0.304±0.426
1/30
PID-SF
30
0.375±0.101
0/30
Table 10: Generalization over 30 unseen random splines with κmax∈[0.021,0.085](\mathrm{m}−1) . All controllers, including NMPC, are evaluated on all 30 tracks.
σobs
RMSEd(\mathrm{m})
0.000
0.314±0.000
0.010
0.308±0.008
0.020
0.369±0.031
0.050
1.549±0.546
0.100
2.015±0.337
Table 11: Observation-noise robustness. Values are mean ± std over five independent noise realizations of the same trained policy; the zero standard deviation at σ=0 reflects the deterministic policy. This differs from the five training seeds of Section 5.7 .
Var.
Description
RMSEd(\mathrm{m})
Enet(\mathrm{k}\mathrm{J})
A
Full (Ph. 0 → 1 → 2 → 3 → 4)
0.379
26.89
B
Phase 1 only
0.332
58.10
C
Phases 1+2 only
0.312
47.51
D
Phase 2 only (no curriculum)
0.323
63.65
Table 12: Curriculum ablation ( 200×103 steps each). Ranking is by Enet since tracking is effectively unaffected, i.e., tracking spread across variants is 0.067(\mathrm{m}) , while energy spread is 36.8(\mathrm{k}\mathrm{J}) .
Controller
RMSEd (m)
En (kJ)
Regen (%)
Status
PPO fine-tuned ( 100 k)
0.273
25.56
48.9
ok
LQR (dynamic model)
0.438
29.30
46.0
ok
PPO zero-shot
0.473
25.00
48.0
ok
NMPC (kinematic model)
0.560
13.69
68.6
ok
PID-SF (kinematic design)
0.740
30.10
44.7
ok
Stanley (re-tuned)
2.008
25.06
52.2
ok
Table 13: Dynamic single-track validation with linear tires, Cf=Crtire=80(\mathrm{k}\mathrm{N},\mathrm{r}\mathrm{a}\mathrm{d}−1),v≤8(\mathrm{m}\mathrm{s}−1) . All runs reach the end of path. NMPC retains its kinematic internal model and operates under plant–model mismatch, a nominal case in practice.
Figure 12: Lateral deviation on the dynamic single-track plant with linear tires and v≤8(\mathrm{m}\mathrm{s}−1) . The fine-tuned policy tracks second only to the purposefully-designed dynamic LQR, while remaining within the linear-tire regime ∣alat∣≈3.48(\mathrm{m}\mathrm{s}−2) .
Appendix figures & tables1 asset
Supplementary material from the paper’s appendix.
Appendix
Phase
wh
wd
wp
wδ
wv
wa
we
0 (longitudinal, straight)
0.3
0.3
1.0
0.02
0.8
0.10
0.0
1 (lateral introduction)
0.8
1.2
0.5
0.05
0.3
0.05
0.0
2 (stopping, full path)
0.8
1.2
0.8
0.05
0.3
0.05
0.0
3 (disturbance recovery)
1.0
1.5
0.5
0.10
0.4
0.05
0.0
4 (energy optimization)
0.6
1.0
0.5
0.10
0.5
0.15
0.3
Appendix
Table 14: Reward weights of ( 22 ) by curriculum phase. An additional penalty of 0.5 is applied whenever the normalized speed error exceeds 0.25 Hess and Ljungbergh (2021) , and rOOB=5 is applied once at the terminal step of a failed episode.