ReBound: Reset-Free Reinforcement Learning for Agile Driving via Reset-Aware Semi-Markov Bootstrapping
Organizations: Department of Mechanical Systems Engineering, Nagoya University, Aichi, Japan · CyberAgent AI Lab, Tokyo, Japan
Abstract
We present ReBound (reset-aware semi-Markov bootstrapping), a value-learning method for reset-free reinforcement learning (RL) that learns agile driving through real-world RL without human intervention after training starts. Reset-free RL explores by alternating between a forward policy that performs the task and a reset policy that restores a drivable state. Standard reset-free RL treats collisions as episode terminations in value-function bootstrapping, truncating the return at each collision. This conflicts with the reset-free objective, under which driving resumes after recovery, and acts as a hidden collision penalty absent from the reward function. ReBound instead applies semi-Markov bootstrapping: it treats the collision-causing action and the subsequent recovery as a single semi-Markov transition and bootstraps from the post-recovery value, discounted by the actual recovery time. We further apply reward centering to stabilize value learning in the resulting continuing task. In simulation and on a 1/10-scale vehicle, ReBound substantially outperforms conventional value-learning methods and model-based control.
Figures & tables
| Environment | Method | Episode Return | Collision Rate |
|---|---|---|---|
| LabTrack-Sim | MPPI | ||
| Terminal-TD-MPC2 | |||
| Timeout-TD-MPC2 | |||
| ReBound-TD-MPC2 | |||
| SMDP only | |||
| RC only |
| Method | Episode Return | Average Speed (m/s) |
|---|---|---|
| Without manual resets (continuing) | ||
| MPPI | ||
| Terminal-TD-MPC2 | ||
| ReBound-TD-MPC2 | ||
| With manual resets (ends at first collision) | ||
| MPPI | ||