Model Predictive Control (MPC) provides a structured and constraint-aware mechanism for decision-making, but its reliance on optimization-friendly analytical dynamics models limits its use in tasks with contacts and other hard-to-model state dependencies. Model-free reinforcement learning avoids explicit modeling assumptions but typically requires large amounts of interaction data. We present a learning-based MPC framework that combines the data efficiency and structure of local model-based planning with learned components that compensate for incomplete dynamics and finite-horizon myopia. The method augments a nominal analytical model with a residual dynamics network that learns missing state-dependent effects from data and combines the resulting planner with a learned action-value critic that injects long-horizon MDP structure into the local iLQR optimization. To make this practical at reinforcement-learning scale, we develop a GPU-accelerated batched iLQR solver that evaluates learned dynamics and critic networks inside the optimal-control loop and solves thousands of trajectory-optimization problems in parallel. The complete system is integrated into a robotics simulator, enabling scalable model-based reinforcement learning under incomplete dynamics. Experiments on biased and incompletely modeled control tasks show that the approach improves closed-loop control performance while preserving the model-based structure needed for efficient constrained trajectory optimization.
Figures & tables
Figure 1: Parallel environments used for training and evaluation. Left: The stack-cube task with a 6-dof manipulator. The robot must pick up the target cube (blue) and stack it on top of the goal cube (red). Right: The quadrotor task, where the drone must hover at a target position while carrying an unmodeled, suspended payload. Both tasks are simulated using Genesis [ 2 ] .
Figure 2: Point-mass reach task with the analytical LQR critic Q∗,V∗ . Columns: success rate and mean episode cost. Top: all-stage blend, λ=λf swept jointly. Bottom: terminal-only blend, λ=1 and λf swept.
Figure 3: Quadrotor learning progress while decreasing the critic blend parameter λ . Early updates behave like nominal MPC and fail under the biased incomplete model (cf. Section 3.1 ); once the learned critic is allowed to shape the objective more strongly, the success rate increases and the cost drops.
Appendix figures & tables12 assets
Supplementary material from the paper’s appendix.
Appendix
Figure 4: Runtime scaling of one batched iLQR iteration with analytical and autodiff derivatives on CPU and GPU. GPU execution becomes decisive once learned models are part of the solve, while autodiff remains close to analytical derivatives; runtime per QP decreases with batching up to about 104 parallel problems before flattening on the tested hardware.
Figure 5: Two closed-loop trajectories under the biased incomplete quadrotor model. Solid lines denote the trajectories stabilized by the learned critic controller; dashed lines denote the nominal MPC. Different random initial states are shown for the two controllers. The nominal controller initially moves toward the goal but excites the suspended load, which then lead to tip-over late into the flight. The learned controller approaches the goal more conservatively, keeping the pendulum motion bounded and stabilizing the system throughout the episode.
Figure 6: Monte Carlo envelopes over N=300 randomized rollouts. Left: blended-critic MPC. Right: nominal MPC. The learned controller keeps the mean close to the hover target and prevents late-episode variance growth. The nominal controller develops large dispersion and a systematic loss of altitude, consistent with repeated crash trajectories.
Figure 7: Local x - y slices of the four learned value heads at one MPC optimization snapshot. All state and action entries except planar position are held fixed. The heads agree on the spatial value distribution, indicating low ensemble uncertainty for the critic correction at this rollout state.
Figure 8: Local critic analysis for the same slice as Figure 7 . The learned value remains consistent with the stage cost by avoiding low-value regions far from the goal, but shifts the local optimum according to the fixed velocity, attitude, and payload state. The critic is smooth in the plotted region, and the value-head variance defines the guidance scale used to blend the critic with the nominal MPC objective.
H
Nominal SR
Nominal cost
Critic SR
Critic cost
2
0.00
–
0.03
278
5
0.116
280
0.685
266
10
0.318
254
0.885
223
20
0.361
251
0.941
220
30
0.298
253
0.018
335
50
0.0
411
0.02
421
Appendix
Table 1: Influence of iLQR horizon length on cube-stacking performance at Δtctrl=0.05s . Nominal MPC requires an intermediate horizon, whereas the learned critic maintains strong performance even for short horizons.
Figure 9: Cube-stacking phases. (1) Approach and Grasp; (2) Pickup and Transport; (3) Align and Drop; Depth (B) and haptic observations (contact-probes in gripper fingers visualized by red markers, figure A1..3) improve grasp estimation. The predicted cube trajectory (C2, blue) matches the predicted end-effector trajectory (C2, green) during the transport phase (2).
Critic input
SR
Mean cost
[x,u,pgoal]
0.230
238.0
[x,u,pgoal,depth]
0.543
234.0
[x,u,pgoal,depth,haptic]
0.941
257.0
Appendix
Table 2: Influence of critic input features on cube-stacking performance. Adding depth and haptic observations improves success rate monotonically, indicating that richer observations improve grasp-phase inference and critic stability.
Experiment
SR
Mean cost
Baseline
0.394
234.7
g˙←[x,u,contact]
0.341
275.6
g˙←[contact,depth]
0.401
278.2
g˙←[x,u,contact,depth]
0.323
273.4
x˙←[x,u,contact]
0.361
296.4
Appendix
Table 3: Influence of residual dynamics learning on cube-stacking performance. Residual variants do not improve significantly over the nominal baseline and often increase cost.
Setting
Value
Planning time step Δtplan
0.05s
Horizon H
20
Parallel environments
100
iLQR iterations
8
Maximum line-search trials
9
Maximum regularization updates
10
Appendix
Table 4: Shared solver, optimization, and data-collection settings used in both tasks.
Setting
Quadrotor
Cube stacking
Episode length
5s
8s
Control time step Δtctrl
0.005s
0.01s
Residual model
ELU MLP, 3 layers, width 128
ELU MLP, 3 layers, width 256
Critic trunk
ELU MLP, 3 layers, width 128
ELU MLP, 3 layers, width 256
Value ensemble
4 heads, 1 layer, width 128 , mean aggregation
4 heads, 1 layer, width 256 , mean aggregation
Appendix
Table 5: Task-specific timing and network settings.
State-of-the-art model-based Reinforcement Learning (RL) approaches either use gradient-free, population-based methods for planning, learned policy networks, or a combination of policy networks and planning. Hybrid approaches that combine Model Predictive Control (MPC) with a learned model and a policy prior to leverage the advantages of both paradigms have shown promising results. However, these approaches typically rely on gradient-free optimization methods, which can be computationally expensive for high-dimensional control tasks. While gradient-based methods are a promising alternative, recent works have empirically shown that gradient-based methods often perform worse than their gradient-free counterparts. We propose Dream-MPC, a novel approach that generates few candidate trajectories from a rolled-out policy and optimizes each trajectory by gradient ascent using a learned world model, uncertainty regularization and amortization of optimization iterations over time by reusing previously optimized actions. Our results on 24 continuous control tasks show that Dream-MPC can significantly improve the performance of the underlying policy and can outperform gradient-free MPC and state-of-the-art baselines. Code and videos are available at https://dream-mpc.github.io.
Jonathan Spieler, Sven Behnke
Autonomous Intelligent Systems, Computer Science Institute VI - Intelligent Systems and Robotics, Center for Robotics and the Lamarr Institute for Machine Learning and Artificial Intelligence, University of Bonn, Germany.
Planning with learned world models combines online trajectory optimization with learned value and policy functions for high-dimensional control. Because the planner determines the experience used for learning, while the learned critic and actor in turn score and propose future plans, planning and learning form a closed feedback loop. TD-MPC is a prominent instance of this design. Recent policy-constrained variants strengthen one part of the loop by aligning the learned policy with planner behavior. We introduce PL-MPC (Planning-Learning MPC), which additionally modifies critic supervision and planner terminal-value estimation. Hybrid multi-step TD targets expose critic updates to more realized rewards before bootstrapping; disagreement-aware terminal estimates reduce the influence of uncertain critic values during MPPI planning; and return-weighted actor distillation emphasizes planner-executed actions from high-return episodes. The world-model architecture and MPPI optimizer are otherwise unchanged. On HumanoidBench, the largest gains occur on \texttt{balance-hard}, where Total Average Return (TAR) increases from 98±18 to 387±255, and \texttt{hurdle}, from 199±13 to 466±200; performance across the broader benchmark remains task dependent, and PL-MPC remains competitive on DMControl. Controlled ablations show different component interactions across the two tasks. We further demonstrate zero-shot sim-to-real transfer on wrench-nut alignment with a 7-DoF KUKA IIWA14, obtaining higher observed success than TD-M(PC)^2 on the training object size and two unseen sizes. Code and data will be available at: https://pl-mpc-humanoid.github.io.
Model Predictive Control (MPC) is widely used in industrial and robotic systems for enforcing constraints and embedding domain knowledge through finite-horizon optimization-based planning. However, despite these strengths, an MPC scheme typically does not yield optimal policies for sequential decision-making problems formulated as Markov Decision Processes (MDPs). Recent combinations of MPC with Reinforcement Learning (RL) alleviate this issue by treating MPC as a parameterized model of the optimal policy of an MDP and adjusting its parameters using data. While these approaches typically consider classical MDPs, many real-world problems include future information--such as forecasts, prices, or reference trajectories--at decision time, which must be included in the MDP state for optimal decision-making. Current MPC-RL approaches do not directly account for this augmented-state structure, raising the question of how to incorporate future information into MPC to obtain an optimal policy. This work establishes the structural requirements under which a parameterized MPC can exactly represent the optimal value functions and policy of an MDP with future information. We further demonstrate that such a parameterized MPC can serve as a structured function approximator, with its parameters learned using RL. The approach is illustrated on a point-mass racing task with future reference information.
Shambhuraj Sawant, Akhil S Anand, Dirk Reinhardt +1
Norwegian University of Science and Technology (NTNU), Trondheim, Norway.