We propose an imitation-learning design for neural-network control policies under state and input constraints. Training alternates a standard imitation gradient step with a block of k safety steps that pull the network's actions toward their projection onto the safe set; at run time, the controller is the trained network alone, with no safety filter. We analyze this scheme as inexact projected gradient descent in the space of policy actions. When the projected actions are recomputed at every safety step and each step moves the actions consistently toward the safe set, letting k grow logarithmically yields asymptotic constraint satisfaction on the training states and bounds the distance to the constrained optimum of the imitation loss; with the projected actions held fixed, the same holds only if they are exactly representable by the network. On a nonlinear autonomous racing task, we compare our method with adding a weighted constraint-violation penalty to the imitation loss. With a sufficiently large weight, our method matches the lap time of unconstrained imitation while reducing the fraction of violating episodes from 15% to 1%, about six times fewer than the penalty approach at its best weight. Its lap times are less sensitive to the weight, which instead sets how quickly violations vanish during training. In racing, the safety corrections are sparse and the conditions of the analysis do not hold; the gain arises instead through the data collected during training. These gains come at the cost of additional training computation.
Figures & tables
Fused
Interleaved
λ
Steps
Dist.
Viol.
c
Steps
Dist.
Viol.
0
12k
4.24
58%
1
1.2k
0.18
0%
10
4k
0.39
49%
2
2.0k
0.15
1%
30
4k
0.16
24%
4
3.8k
0.14
0%
100
4k
0.11
10%
300
4k
0.13
2%
TABLE I: Example 1 , mean over three seeds: gradient steps, RMS distance to v⋆ , and closed-loop episodes leaving the true bounds. Fused with λ=0 minimizes the cost alone; interleaving is evaluated after 200 iterations.
Method
Safety updates
Violations
Lap time
Plain BC on raw demos
–
55/360
5.144±.002
Plain BC on filtered demos
–
24/360
5.189±.007
Fused
1
149/720
5.209±.023
Inter ( K=1 )
1
84/720
5.142±.006
Fused (compute matched)
Kt
41/720
5.192±.014
Inter (log., fixed)
Kt
44/720
5.142±.006
TABLE II: Ablations at λ=10 with 120 starts per seed: six seeds for the middle block and three for the others. The references use no safety loss. Lap time is mean ± SD of seed-level means (s). All methods refresh targets except Inter (log., fixed).
Fig. 1: Closed-loop trajectories on the 40-start illustration bank, with six learned-policy seeds at λ=10 and one run per expert configuration. Gray lines mark the track boundaries; each episode is drawn until its first violation, with the final 2.5 s before it in red and the violation marked by a circle. Labels give the training update, mean lap time, and violating episodes, where Uη,β(θ)=θ−η∇J(θ)−β∇ΦC(θ) is the fused step, T=300 is the number of epochs, and Mt is the number of minibatches in epoch t . This bank shares two starts with the primary 120-start bank.
Fig. 2: Violating episodes during training (mean over seeds; six seeds at λ=10 , three at λ=100 ), evaluated without a filter on ten starts every ten epochs. Before epoch 90, nearly every episode violates in all configurations.
Fused Gradient
Interleaving (proposed)
λ
Violations
Lap time
Violations
Lap time
0.01
10/7/18
5.149±.010
8/9/12
5.139±.004
0.1
10/3/8
5.136±.007
20/3/27
5.146±.010
1
33/2/7
5.159±.006
8/20/30
5.151±.012
10
8/25/12/73/18/13
5.209±.023
0/0/0/0/5/2
5.139±.006
100
28/27/13
5.203±.011
0/3/3
5.137±.002
TABLE III: Safety-weight sensitivity: violating episodes per seed (out of 120 each) and lap time, mean ± SD of seed-level means (s). The λ=100 runs and seeds four to six at λ=10 were added after the initial sweep.
Fig. 3: Lap time and violation rate across safety weights. Bands show SD of seed-level lap means; dots show individual seed violation rates and squares their means. The shared dotted line marks the expert’s 5.171 s lap time and 25.8% violation rate on the respective axes.
Imitation learning (IL) is an effective approach to train complex robotics policies. Recent works have introduced hard constraints into imitation-learning optimization problems to ensure safety, stability, and robustness of the learned policy. However, we argue that these constraints are sometimes infeasible, which can lead to unstable or difficult training dynamics. We study a simple remedy for such situations based on recent theoretical results on the augmented Lagrangian method in infeasible settings. We show that our approach drives the learned policy toward the solution of a closest-feasible constrained IL problem with desirable properties. The method is illustrated on a toy driving example with a total-acceleration constraint and pedestrian-safety constraints, a setting in which infeasibility can naturally arise while still allowing a safe learned policy.
Recent imitation learning (IL) algorithms such as flow-matching and diffusion policies demonstrate remarkable performance in learning complex manipulation tasks. However, these policies often fail even when operating within their training distribution due to extreme sensitivity to initial conditions and irreducible approximation errors that lead to compounding drift. This makes it unsafe to deploy IL policies in the field where out-of-distribution scenarios are prevalent. A prerequisite for safe deployment is enabling the policy to determine whether it can execute a task the way it was learned from demonstrations. This paper presents TAIL-Safe, a principled approach to identify, for a trained IL policy, a safe set from where the policy empirically succeeds in completing the learned task. We propose a Lipschitz-continuous Q-value function that maps state-action pairs to a long-term safety score based on three short-term task-agnostic criteria: visibility, recognizability, and graspability. The zero-superlevel set of this function characterizes an empirical control invariant set over state-action pairs. When the nominal policy proposes an action outside this set, we apply a recovery mechanism inspired by Nagumo's theorem that uses gradient ascent to the Q-function to steer the policy back to safety. To learn this Q-function, we construct a high-fidelity digital twin using Gaussian Splatting that enables systematic collection of failure data without risk to physical hardware. Experiments with a Franka Emika robot demonstrate that flow-matching policies, which fail under run-time perturbations, achieve consistent task success when guided by the proposed TAIL-Safe.
While neural network control policies are powerful, their deployment on safety critical systems depends on ensuring that they obey strict constraints. Existing work often treats safety as a metric to optimize for, which competes with other performance objectives, if training converges at all. Instead, we introduce ShardNet, a neural network architecture that strictly enforces unions of polyhedral constraints by construction, using a differentiable projection layer parameterized by a classification network. The key insight is to embed safety into the neural network's structure, allowing performance to be optimized independently because formal safety guarantees are always given. In contrast with existing neural architectures that can only enforce simple convex constraints, ShardNet enables the first safe-by-construction synthesis of forward-invariant neural network controllers on closed-loop systems where safety constraints are expressed as nonconvex unions of polyhedras or learned value function level sets. To support this, we also introduce a technique to verify and train such value functions correctly as rectified linear unit (ReLU) networks, which has not previously been possible. On double integrator benchmarks drawn from the literature, ShardNet policies maintain 100% safety on verified sets and achieves significantly lower objective loss compared to existing formal methods. Furthermore, our value function training technique also produces safe sets more than 3 times larger than existing verification approaches.
Long Kiu Chung, Shreyas Kousik
Department of Mechanical Engineering, Georgia Institute of Technology, Atlanta, GA.