Real-time whole-body controllers for legged robots typically plan through a fixed nominal model and degrade when the deployed dynamics change. Adaptive methods typically require a model structure that contact dynamics do not provide, or they need offline training for each anticipated condition. We present Look-back and Look-ahead Adaptive Model Predictive Path Integral control (LLA-MPPI). The method converts whole-body adaptation into selection over a bank of GPU-batched contact simulators with different physical or structural parameters. Windowed prediction errors select the simulator that best explains recent motion. A whole-body MPPI planner optimizes controls through the selected model. The framework requires no offline training, and its selected hypotheses are physically interpretable. Across four simulated tasks, it achieves 97.5% success while the strongest baseline reaches 74% and an oracle with the true model reaches 98.5%. Hardware validation on a Unitree Go2 shows the robot walking under a payload added mid-run, walking after one leg is disabled, and pushing a box to its goal while increasing its mass on the fly. Code, videos, and project details are available at: https://lla-control.github.io
Figures & tables
Fig. 2: Overview of the LLA-MPPI framework. The GPU-parallel look-back stage selects the model that best explains recent state transitions, while the CPU-parallel look-ahead stage uses the selected model to solve the MPPI. The two stages operate asynchronously to provide real-time adaptive control.
Fig. 3: Look-back update time versus bank size K for the GPU-batched MJWarp bank and a CPU MuJoCo baseline. Markers and error bars denote mean and spread across runs. Performed on an NVIDIA RTX 4070 Laptop GPU (8 GB) and an AMD Ryzen 9 8945HS CPU (8 cores / 16 threads).
Fig. 4: Success rates over 50 randomized trials per task. Error bars denote 95% Wilson confidence intervals. Success is defined as reaching the task goal within the episode.
Fig. 5: Time-to-goal on successful trials only (failed runs are excluded). Boxes show the interquartile range (IQR), red horizontal lines indicate medians, whiskers extend to the most extreme points within 1.5× IQR. Sample size n is the number of successful trials out of 50.
Method
Success rate (%)
Tasks ≥90%
Fall rate (%)
Solve time (ms)
Explicit ID
MPPI
25.5 [ 20.0 , 32.0 ]
1/4
65.3 [ 57.4 , 72.5 ]
16.57
—
Oracle-MPPI
98.5 [ 95.7 , 99.5 ]
4/4
0.0 [ 0.0 , 2.5 ]
17.77
true model (given)
LLA-MPPI
97.5 [ 94.3 , 98.9 ]
4/4
0.0 [ 0.0 , 2.5 ]
17.60
discrete bank
DOB
42.5 [ 35.9 , 49.4 ]
1/4
33.3 [ 26.3 , 41.2 ]
15.43
—
CPE
74.0 [ 67.5 , 79.6 ]
1/4
21.3 [ 15.5 , 28.6 ]
18.54
continuous θ
TABLE I: Aggregate results over four tasks with 50 trials each (200 episodes per method). All trials were run on a single workstation (Intel Core i9-9980XE, 18 cores / 36 threads; NVIDIA RTX 2080 Ti, 11 GB; Ubuntu 22.04; MuJoCo 3.4.0 with MuJoCo-Warp 1.12.0), using one GPU per run. Brackets denote 95% Wilson confidence intervals. Fall rates exclude the leg-lock task, where falls are not defined by the protocol. Solve time is the mean synchronous controller-update latency, and Explicit ID indicates whether a method produces a physical parameter estimate. Bold marks the best value per column among the realizable methods, and the gray row is the oracle reference.
Fig. 6: Windowed one-step prediction error during the asymmetric-payload task. The shaded region spans the minimum-to-maximum error across the model bank, and dashed lines indicate payload changes.
Fig. 7: Asymmetric-payload results. Top: mean planar trajectories relative to the reference path, with parentheses indicating successful trials out of 50 and dotted curves denoting mean failed trajectories. Bottom: estimated left and right payload masses from LLA-MPPI and CPE.
Fig. 8: Results for the slope and friction task. Top: mean planar trajectories. Bottom: friction and terrain-slope estimates relative to ground truth. LLA-MPPI and CPE track both changes and complete all trials, whereas nominal MPPI and DOB largely fail.
Fig. 9: Ablation of lookback window W and bank size K on the slope-friction task with noisy state measurements. Results show mean friction-estimation error over 10 paired seeds, with 95% confidence intervals. The star marks the minimum over the tested settings at (K,W)=(257,25) .
Fig. 10: Representative trajectories for the asymmetric-payload RL comparison. Markers show the 3.5kg right-side payload addition, its transfer to the left, endpoints, and falls. LLA-MPPI and both PPO-DR policies reach the goal, whereas nominal MPPI and PPO fall.
Method
Offline
90%
95%
99%
Online
PPO-DR (nominal)
48.5 min
46 M
46 M
69 M
<1 ms
PPO-DR (heavy)
54.7 min
92 M
115 M
161 M
<1 ms
LLA-MPPI
–
–
–
–
14.4±1.2 ms
TABLE II: Training and online computational cost for the asymmetric payload comparison. The percentage columns report the first PPO checkpoint reaching the indicated fraction of its final mean velocity-tracking return. Offline denotes PPO training wall time on an RTX 4070 Laptop GPU; online denotes computation per control step, including MPPI and model-bank computation for LLA-MPPI .
Humanoid robots require whole-body controllers that are both robust and precise in contact-rich environments. While deep reinforcement learning (RL) achieves robust stability, its behavior is tightly coupled to the training objective and command interface, making it difficult to add new feedback objectives without retraining. In this study, we propose an RL guided whole-body model predictive path integral (MPPI) framework that acts as an add-on feedback controller on top of a pretrained RL policy. Instead of using RL policy as the final controller, we use it as a sampling prior that biases MPPI rollouts toward dynamically feasible behaviors. Task objectives are specified through modular MPPI cost terms, and MPPI closes the loop by continuously correcting the RL prior online to satisfy these objectives without retraining the policy. Simulations on a 29-DoF Unitree G1 humanoid in MuJoCo demonstrate stable high-rate control (average 280~Hz). The proposed method improves task-level precision over a pure RL baseline under the same command interface. This is achieved by correcting systematic drift during straight walking and tracking additional whole-body reference signals imposed through the cost.
Yunsoo Seo, Sol Choi, Euncheol Im +2
Center for Humanoid Research, Korea Institute of Science and Technology (KIST), 02792 Seoul, South Korea · University of Texas at Austin, Austin, Texas 78712, USA · Department of Mechanical Engineering, Yonsei University, 03722, Seoul, South Korea +1
This paper presents a multi-phase whole-body model predictive control (MPC) approach for bipedal walking, combining a detailed whole-body model in the near horizon with a simplified single-rigid-body model in the later prediction steps. This reduces computational complexity while retaining prediction capabilities. The resulting nonlinear optimal control problem is solved entirely within the general-purpose, off-the-shelf nonlinear MPC framework acados, using sequential quadratic programming (SQP). Given a contact schedule and a target walking speed, the controller optimizes joint torques without depending on preselected footstep locations. The controller is validated in MuJoCo simulation on the 18-DoF bipedal robot HyPer-2.
Franek Stark, Felix Wiebe, Shubham Vyas +2
Robotics Innovation Center at the German Research Center for Artificial Intelligence (DFKI), Bremen, Germany · University Bremen, Germany
This work proposes the L1 Adaptive Model Predictive Path Integral (L1-MPPI). It cascades L1 adaptive control with the Model Predictive Path Integral (MPPI) to improve tracking of high-speed UAV trajectories. Thanks to the L1augmentation, the tracking remains accurate even under model uncertainties and external disturbances, such as an additional payload or a mismatch in the modeled aerodynamic drag. In contrast to existing MPPI approaches for UAV control that do not explicitly model aerodynamic effects, varying payloads, and typically neglect the dynamics of low-level motor controllers, our L1-MPPI approach enhances the dynamic model used in the MPPI by incorporating the low-level flight controller and motor dynamics, as well as an iterative mixing scheme that reflects the approach of the low-level controller. The proposed method demonstrates improved tracking performance in both simulation and the real world, even when the UAV is subjected to an unknown payload. In flight with 35% mass increase, our approach lowers the RMSE by 58.61% with respect to plain MPPI. Compared to the same MPPI using an online mass estimator in place of the L1 augmentation, the RMSE is lower by 38.59%. During the real-world experiments the UAV reaches speeds up to 13.50 m/s and accelerations up to 2.5 g.
Lukáš Kotek, Ondřej Procházka, Vojtěch Vonásek +2
Multi-robot Systems Group, Faculty of Electrical Engineering, Czech Technical University in Prague, Czech Republic