We present ReBound (reset-aware semi-Markov bootstrapping), a value-learning method for reset-free reinforcement learning (RL) that learns agile driving through real-world RL without human intervention after training starts. Reset-free RL explores by alternating between a forward policy that performs the task and a reset policy that restores a drivable state. Standard reset-free RL treats collisions as episode terminations in value-function bootstrapping, truncating the return at each collision. This conflicts with the reset-free objective, under which driving resumes after recovery, and acts as a hidden collision penalty absent from the reward function. ReBound instead applies semi-Markov bootstrapping: it treats the collision-causing action and the subsequent recovery as a single semi-Markov transition and bootstraps from the post-recovery value, discounted by the actual recovery time. We further apply reward centering to stabilize value learning in the resulting continuing task. In simulation and on a 1/10-scale vehicle, ReBound substantially outperforms conventional value-learning methods and model-based control.
Figures & tables
Fig. 2 : Learning curves on each simulated course. Lines show the mean, and bands indicate the standard deviation; dashed lines show the average evaluation return of MPPI.
Fig. 3 : Driving trajectories during evaluation on each simulated course (5 seeds × 10 episodes overlaid). Colors indicate speed, black triangles indicate start positions, and red × symbols mark collisions. ReBound-TD-MPC2 approaches the speed limit on straights and decelerates before corners, adapting its driving to each course.
Environment
Method
Episode Return
Collision Rate
LabTrack-Sim
MPPI
7.6±3.9
0.00
Terminal-TD-MPC2
52.9±64.4
0.70
Timeout-TD-MPC2
11.1±4.5
1.00
ReBound-TD-MPC2
139.7±10.4
0.03
SMDP only
47.4±59.6
0.50
RC only
18.2±10.3
0.23
TABLE I : Simulation results (mean ± standard deviation over 5 seeds × 10 episodes). Italics denote ablations of ReBound-TD-MPC2.
Fig. 4 : Value function analysis on LabTrack-Sim . (a) Solid lines: Q-values of each critic; dashed lines: learning targets. Qθterm flattens the sharp drop in Gterm before a collision, whereas QθReBound follows G . (b) Mean ± standard error over 3 seeds; the dash-dotted line marks the observed average recovery length (38 steps). ReBound retains a larger fraction of the post-collision value as recovery becomes faster, whereas Terminal does not.
Fig. 5 : Driving after training on the real LabTrack with the reset policy enabled. Markers distinguish episodes, and numbers indicate start positions.
Fig. 6 : Episode returns during training on the real LabTrack . Moving average over 50 episodes; bands show the standard error within the window. The dashed line and gray band show the mean and standard deviation of MPPI’s evaluation returns.
Method
Episode Return
Average Speed (m/s)
Without manual resets (continuing)
MPPI
1.67±2.74
0.40±0.60
Terminal-TD-MPC2
3.89±3.62
1.36±0.55
ReBound-TD-MPC2
68.26±56.12
2.00±0.66
With manual resets (ends at first collision)
MPPI
8.62±3.83
1.56±0.38
TABLE II : Evaluation results over 10 episodes on the real LabTrack (mean ± standard deviation). Without manual resets: the reset policy recovers the vehicle, and driving continues for 500 steps. With manual resets: a person positions the vehicle, and the episode ends at the first collision.
Should a single collision necessarily terminate an entire navigation episode? In most deep reinforcement learning (DRL) frameworks for robot navigation, this remains the standard practice: every collision immediately triggers a global environment reset and is penalized as a complete task failure. While a collision during deployment naturally indicates task failure, applying the same treatment during training prevents the agent from exploring challenging obstacle configurations, which slows learning progress in the early training phase. In this work, we challenge this convention and propose a Multi-Collision reset Budget (MCB) framework that decouples local collision termination from global environment resets, allowing the agent to retry difficult configurations within the same episode. Simulation experiments show that MCB improves early-stage learning efficiency by reaching target success-rate levels with fewer interactions, with small collision budgets producing the most consistent gains. Real-world experiments on heterogeneous robot platforms further validate the deployability of the learned policies in cluttered environments.
Shanze Wang, Xinming Zhang, Siwei Cheng +4
College of Information Science and Technology, Eastern Institute of Technology, Ningbo, China · Department of Aeronautical and Aviation Engineering, The Hong Kong Polytechnic University, Hung Hom, Kowloon, Hong Kong · School of Computer Science and Technology, University of Science and Technology of China, Hefei, China +2
Traditional reinforcement learning (RL) for recovery in autonomous systems lacks causal understanding and generalizes poorly to novel failure scenarios. RL policies often stall in failure states, spending up to 70% of an episode immobilized. Rule-based recovery alone is inadequate, and adding heuristic recovery to a pretrained PPO policy worsens rewards because policies cannot coordinate well with unanticipated interventions. The issue is not missing recovery mechanisms but a lack of policies trained to collaborate with them. We introduce CRRL, a causal-guided RL framework that trains policies to work effectively with rule-based recovery. The recovery detects stalled states and assists the agent. Causal relations from driving logs shape the training signal, teaching the policy to anticipate stalls and adjust actions in recovery contexts. The framework follows MAPE-K, with sensor collection, causal model construction, and hybrid RL policy training corresponding to Monitor, Analyze, and Plan/Execute, respectively. We evaluate CRRL through a four-condition ablation study across three driving scenarios, with 20 episodes per condition. We find that causal training significantly improves reward, distance, and velocity. Moreover, 9 of 20 roundabout episodes required zero recovery intervention, confirming navigation competence. These results show that causal-guided training produces effective RL policies that cooperate with rule-based safety components.
Safia Fatima, Kai Olav Ellefsen, Leon Moonen
Simula Research Laboratory, Norway · University of Oslo, Norway
Safe reinforcement learning seeks policies that maximise task performance while satisfying safety constraints. In driving benchmarks, however, collision costs typically appear only at the time of collision, providing no advance warning of an approaching hazard. Frozen vision--language models can provide dense semantic feedback, yet it remains unclear whether their scores anticipate collisions and which component drives an observed safety improvement. Episodic cost can also favour policies that make little task progress. To address these gaps, we propose VLM-Safe-RL, a framework that integrates frozen CLIP signals into PPO-Lagrangian through reward shaping and an augmented multiplier update. On MetaDrive Hard, which combines the densest traffic with the largest map, the catastrophe rate falls from 31.6% to 19.4%. FormulaOne-L2 analysis finds no evidence that the CLIP signals anticipate collisions and shows that the VLM term has a negligible effect on the Lagrange multiplier. These findings show a conditional reduction in observed catastrophe rate without evidence of collision anticipation.