We propose an imitation-learning design for neural-network control policies under state and input constraints. Training alternates a standard imitation gradient step with a block of k safety steps that pull the network's actions toward their projection onto the safe set; at run time, the controller is the trained network alone, with no safety filter. We analyze this scheme as inexact projected gradient descent in the space of policy actions. When the projected actions are recomputed at every safety step and each step moves the actions consistently toward the safe set, letting k grow logarithmically yields asymptotic constraint satisfaction on the training states and bounds the distance to the constrained optimum of the imitation loss; with the projected actions held fixed, the same holds only if they are exactly representable by the network. On a nonlinear autonomous racing task, we compare our method with adding a weighted constraint-violation penalty to the imitation loss. With a sufficiently large weight, our method matches the lap time of unconstrained imitation while reducing the fraction of violating episodes from 15% to 1%, about six times fewer than the penalty approach at its best weight. Its lap times are less sensitive to the weight, which instead sets how quickly violations vanish during training. In racing, the safety corrections are sparse and the conditions of the analysis do not hold; the gain arises instead through the data collected during training. These gains come at the cost of additional training computation.
Figures & tables
Fused
Interleaved
λ
Steps
Dist.
Viol.
c
Steps
Dist.
Viol.
0
12k
4.24
58%
1
1.2k
0.18
0%
10
4k
0.39
49%
2
2.0k
0.15
1%
30
4k
0.16
24%
4
3.8k
0.14
0%
100
4k
0.11
10%
300
4k
0.13
2%
TABLE I: Example 1 , mean over three seeds: gradient steps, RMS distance to v⋆ , and closed-loop episodes leaving the true bounds. Fused with λ=0 minimizes the cost alone; interleaving is evaluated after 200 iterations.
Method
Safety updates
Violations
Lap time
Plain BC on raw demos
–
55/360
5.144±.002
Plain BC on filtered demos
–
24/360
5.189±.007
Fused
1
149/720
5.209±.023
Inter ( K=1 )
1
84/720
5.142±.006
Fused (compute matched)
Kt
41/720
5.192±.014
Inter (log., fixed)
Kt
44/720
5.142±.006
TABLE II: Ablations at λ=10 with 120 starts per seed: six seeds for the middle block and three for the others. The references use no safety loss. Lap time is mean ± SD of seed-level means (s). All methods refresh targets except Inter (log., fixed).
Fig. 1: Closed-loop trajectories on the 40-start illustration bank, with six learned-policy seeds at λ=10 and one run per expert configuration. Gray lines mark the track boundaries; each episode is drawn until its first violation, with the final 2.5 s before it in red and the violation marked by a circle. Labels give the training update, mean lap time, and violating episodes, where Uη,β(θ)=θ−η∇J(θ)−β∇ΦC(θ) is the fused step, T=300 is the number of epochs, and Mt is the number of minibatches in epoch t . This bank shares two starts with the primary 120-start bank.
Fig. 2: Violating episodes during training (mean over seeds; six seeds at λ=10 , three at λ=100 ), evaluated without a filter on ten starts every ten epochs. Before epoch 90, nearly every episode violates in all configurations.
Fused Gradient
Interleaving (proposed)
λ
Violations
Lap time
Violations
Lap time
0.01
10/7/18
5.149±.010
8/9/12
5.139±.004
0.1
10/3/8
5.136±.007
20/3/27
5.146±.010
1
33/2/7
5.159±.006
8/20/30
5.151±.012
10
8/25/12/73/18/13
5.209±.023
0/0/0/0/5/2
5.139±.006
100
28/27/13
5.203±.011
0/3/3
5.137±.002
TABLE III: Safety-weight sensitivity: violating episodes per seed (out of 120 each) and lap time, mean ± SD of seed-level means (s). The λ=100 runs and seeds four to six at λ=10 were added after the initial sweep.
Fig. 3: Lap time and violation rate across safety weights. Bands show SD of seed-level lap means; dots show individual seed violation rates and squares their means. The shared dotted line marks the expert’s 5.171 s lap time and 25.8% violation rate on the respective axes.