We present RACER, a hierarchical control framework for wheel-based quadruped racing that combines an MPPI planner with a learned residual dynamics model and a low-level RL velocity tracker. The planner augments a nominal unicycle kinematic model with a neural residual term to capture the closed-loop tracking behavior of the RL policy. To train this residual model under limited real-world data, we propose Low-Rank Residual Adaptation (LoRRA), a two-stage approach that pre-trains on large-scale simulation data for broad coverage and then fine-tunes on a small real-world dataset with a low-rank constraint. In simulation, we empirically validate our engineering choices by showing (A) Residual dynamics improve the overall performance of our pipeline by capturing the tracking error of RL velocity tracker at high-speed cornering. (B) Residual dynamics trained with both source-domain and target-domain data gives racing performance significantly better than the residual dynamics trained with only target-domain data. (C) Low-rank constraint at target-domain adaptation gives higher success rates and higher performance than full-tune and from-scratch when domain gap in ground coefficient or joint gain increases.
Figures & tables
Fig. 1 : A Go2W Wheeled Quadruped racing along L-shape track using RACER framework. The low-level controller is a velocity RL-tracking policy; the high level planner is a DIAL-MPC controller equipped with residual dynamics trained with LoRRA.
Fig. 2 : System overview. An DIAL-MPC planner samples velocity commands at 10 Hz and rolls them out under f=funi+gϕ∗ . An RL tracker realises the chosen command at 50 Hz, producing 16-D joint targets for the Go2-W’s 12 leg joints and 4 wheel motors. Hardware adaptation refines only gϕ∗ to compensate domain shift, leaving the planner and tracker untouched.
Fig. 3 : The Source-Domain loss of residual dynamics L(ϕ,Ds) through the training with target-domain data Dt . See Sec. IV-B for definition of each training recipe: LoRRA, Full Tune, and From Scratch. Source and target environments are defined to be two simulation dynamics with different tire frictions, ∣μ∣=0.4 . See Sec. IV-B1 for more details about domain definition. LoRRA(dual-phase training with low-rank constraint) helps preserve the most source-domain posterior with lowest source-domain loss after target-domain adaptation.
Fig. 4 : Simulation performance of Naive Unicycle (Blue) and RACER (Red) v.s. Capped velocity. RACER has stable performances independent of the velocity cap while the performance of naive unicycle is downgraded as the cap becomes higher. Horizontal axis: cap velocity. Vertical Axis: the performance metrics. See Sec. IV-A for the definition of naive unicycle. See Sec. IV-A1 for interpretations.
Fig. 5 : Residual Magnitude under different linear velocities(horizontal axis) and yaw rate changes(vertical axis). Residual magnitude positively correlates with linear velocities and yaw rate changes. Left: Residual of linear velocity; Right: residual of angular velocity. The brighter the area, the larger the residual.
Fig. 6 : Next-State Prediction Errors of RACER(left) and Naive Unicycle(right). Naive Unicycle’s prediction errors explode while cornering while RACER keeps the next-state prediction errors consistenly low. Heat Bar: next-state prediction errors. The brighter the points, the larger the prediction errors. Left: ∥st+1−(funi(st)+gϕ(st))∥ . Right: ∥st+1−funi(st)∥ . The trajectories of left are collected using RACER; The trajectories of right are collected using naive unicycle. A cross dot represents a lap failure.
velocity cap
Method
Clean laps
Avg. vel.
Max dev.
(m/s)
(of 3)
(m/s)
(m)
1.5
RACER
3.0±0.0
1.21±0.01
0.42±0.13
Naive Unicycle
2.3±1.2
1.49±0.04
0.59±0.12
2.0
RACER
3.0±0.0
1.41±0.02
0.45±0.08
Naive Unicycle
0.0±0.0
0.00
0.83±0.57
3.2
RACER
3.0±0.0
1.95±0.04
0.35±0.03
TABLE I : Hardware counterpart of Fig. 4 . Racing metrics (mean ± std over three runs) on the physical robot at different velocity caps vx . RACER completes every lap and scales its speed as the caps become higher, whereas the naive unicycle finishes laps only at the conservative 1.5 m/s cap and fails entirely ( 0 laps) at higher speeds. Clean laps are out of 3 per run; bold marks the better value per metric.
Fig. 7 : Performance(Vertical Axis) of RACER using residual dynamics trained with LoRRA(blue), From Scratch(Yellow), Full-Tune(Green) under various domain gaps in terms of friction coefficient(horizontal Axis). The larger the domain gaps, the greater the advantages hold by LoRRA. Shaded Area: variance across seeds. For the definition of each training recipe, see Sec. IV-B
Fig. 8 : The advantage of LoRRA at preserving source-domain posterior increases with domain gap magnitude. The horizontal axis shows the domain gap Δμ . The vertical axis shows the difference in source-domain loss after adapting residual dynamics to Dt , comparing Full-Tune against LoRRA: L(ϕ,Ds)Full-Tune−L(ϕ,Ds)LoRRA .
Fig. 9 : Hardware ablation experiments for RACER and LoRRA.
Fig. 10 : LoRRA gives performance consistently greater than or equal to From Scratch and Full Tune under different joint-gain domain gaps . This figure is the counterpart of Fig. 7 with domain gaps redefined as joint gain difference instead of friction difference.
ΔKp
Method
Clean laps
Avg. vel.
(gap)
(max is 3)
(m/s)
0
LoRRA
3.00±0.00
1.94±0.07
Full-Tune
2.67±0.47
2.06±0.05
From-Scratch
0.33±0.47
0.59±0.83
30
LoRRA
2.33±0.94
1.61±0.08
Full-Tune
1.00±0.82
0.60±0.44
TABLE II : Hardware LoRRA ablation (counterpart of Fig. 10 ). Racing metrics (mean ± std over three independently trained seeds) for the three residual-training recipes under an increasing joint-gain domain gap ΔKp=∣Kp,s−Kp,t∣ (larger ΔKp= larger gap; the residual is pre-trained at the nominal gain Kp,s=70 ). LoRRA completes laps robustly across all gaps, whereas Full-Tune degrades as the gap grows and fails entirely at ΔKp=50 , and From-Scratch fails at nearly every gap. Velocity is averaged over all seeds (a failed run contributes 0 velocity); bold marks the best value per metric.
We present a modular framework to benchmark new and existing methods for trajectory planning and control in high-acceleration maneuvers that push autonomous driving to the limits. Our framework includes time-optimal raceline generation, online time-optimal velocity replanning, geometric path tracking controllers, and a new model-structured neural network (MS-NN) to learn the inverse dynamics for steering control. We deploy our framework on a 1:10-scale RoboRacer platform, using two circuits. Through several ablations with cautious and aggressive racelines, we study the performance of single modules and their combinations. We show that our MS-NN significantly improves tracking accuracy, decreases steering oscillations, and is physically interpretable. Moreover, online velocity replanning improves lap times by compensating for execution errors, and enables the vehicle to safely reach higher speeds and accelerations. To support future research, our code, datasets, videos and results are publicly available at https://roboracer-benchmark.github.io/planning_control_benchmark/.
Mattia Piccinini, Patrick Zambiasi, Aniello Mungiello +3
Professorship of Autonomous Vehicle Systems, Technical University of Munich, 85748 Garching, Germany; Munich Institute of Robotics and Machine Intelligence (MIRMI). · Avilus GmbH, Germany. · Department of Information Technology and Electrical Engineering (DIETI), University of Naples Federico II, Naples 80125, Italy. +1
High-precision locomotion combines motion-command tracking with precise regulation of task-relevant physical states, enabling robots to interact reliably with their surroundings during motion. Joint end-to-end optimization can leave precision objectives insufficiently optimized, while reactive residual control adjusts actions only after deviations become observable. We present \textbf{LocoWM}, a world-model-guided preactive residual adaptation framework for high-precision locomotion. A base policy provides command-following locomotion, while an action-conditioned world model predicts a sequence of future physical states from proprioceptive history and the proposed base action. A residual adapter conditions on this predicted sequence to generate additive action corrections that compensate for anticipated deviations. Two-stage training first learns locomotion and action-conditioned dynamics, then freezes both modules while training the adapter, separating locomotion acquisition from precision adaptation. Experiments spanning terrain leveling, acceleration compensation, and push recovery demonstrate improved control precision and disturbance robustness over end-to-end and reactive residual baselines. Demos and code are available at: https://zhaozijie2022.github.io/LocoWM
Zijie Zhao, Shengqian Chen, Xiaoxu Wang +3
University of Chinese Academy of Sciences · Institute of Automation, Chinese Academy of Sciences · Beijing University of Posts and Telecommunications +1
This paper presents a hierarchical control framework using model predictive control (MPC) and reinforcement learning (RL) for active roll control to manage lateral load transfer during autonomous racing of a wheeled quadruped. The framework integrates offline time-optimal raceline generation, an online MPC planner that actively minimizes the lateral Load Transfer Ratio (LTR), and a low-level, whole-body RL policy deployed directly onto the robot's 16 actuators. The MPC is based on a vehicle dynamics bicycle model of the Unitree Go2-W platform. The robot's leg actuators act as active suspension where knee joints generate anti-roll torque to bank into turns. Physical track experiments demonstrate that active roll control reduces mean LTR by up to 44%, improves the fastest lap time by 8.7%, and boosts peak lateral acceleration capability by 21.3% to 1.98 m/s2, maintaining robust high-speed stability beyond the range of a non-tilting baseline controller. Supplementary code and video can be found at https://github.com/meisman-ucb/go2w-roll-control-mpc
Marla Eisman, Brian Lam, Samuel Sonnino +1
1) University of California, Berkeley · 2) Politecnico di Milano, Italy