Model Predictive Control (MPC) provides a structured and constraint-aware mechanism for decision-making, but its reliance on optimization-friendly analytical dynamics models limits its use in tasks with contacts and other hard-to-model state dependencies. Model-free reinforcement learning avoids explicit modeling assumptions but typically requires large amounts of interaction data. We present a learning-based MPC framework that combines the data efficiency and structure of local model-based planning with learned components that compensate for incomplete dynamics and finite-horizon myopia. The method augments a nominal analytical model with a residual dynamics network that learns missing state-dependent effects from data and combines the resulting planner with a learned action-value critic that injects long-horizon MDP structure into the local iLQR optimization. To make this practical at reinforcement-learning scale, we develop a GPU-accelerated batched iLQR solver that evaluates learned dynamics and critic networks inside the optimal-control loop and solves thousands of trajectory-optimization problems in parallel. The complete system is integrated into a robotics simulator, enabling scalable model-based reinforcement learning under incomplete dynamics. Experiments on biased and incompletely modeled control tasks show that the approach improves closed-loop control performance while preserving the model-based structure needed for efficient constrained trajectory optimization.
Figures & tables
Figure 1: Parallel environments used for training and evaluation. Left: The stack-cube task with a 6-dof manipulator. The robot must pick up the target cube (blue) and stack it on top of the goal cube (red). Right: The quadrotor task, where the drone must hover at a target position while carrying an unmodeled, suspended payload. Both tasks are simulated using Genesis [ 2 ] .
Figure 2: Point-mass reach task with the analytical LQR critic Q∗,V∗ . Columns: success rate and mean episode cost. Top: all-stage blend, λ=λf swept jointly. Bottom: terminal-only blend, λ=1 and λf swept.
Figure 3: Quadrotor learning progress while decreasing the critic blend parameter λ . Early updates behave like nominal MPC and fail under the biased incomplete model (cf. Section 3.1 ); once the learned critic is allowed to shape the objective more strongly, the success rate increases and the cost drops.
Appendix figures & tables12 assets
Supplementary material from the paper’s appendix.
Appendix
Figure 4: Runtime scaling of one batched iLQR iteration with analytical and autodiff derivatives on CPU and GPU. GPU execution becomes decisive once learned models are part of the solve, while autodiff remains close to analytical derivatives; runtime per QP decreases with batching up to about 104 parallel problems before flattening on the tested hardware.
Figure 5: Two closed-loop trajectories under the biased incomplete quadrotor model. Solid lines denote the trajectories stabilized by the learned critic controller; dashed lines denote the nominal MPC. Different random initial states are shown for the two controllers. The nominal controller initially moves toward the goal but excites the suspended load, which then lead to tip-over late into the flight. The learned controller approaches the goal more conservatively, keeping the pendulum motion bounded and stabilizing the system throughout the episode.
Figure 6: Monte Carlo envelopes over N=300 randomized rollouts. Left: blended-critic MPC. Right: nominal MPC. The learned controller keeps the mean close to the hover target and prevents late-episode variance growth. The nominal controller develops large dispersion and a systematic loss of altitude, consistent with repeated crash trajectories.
Figure 7: Local x - y slices of the four learned value heads at one MPC optimization snapshot. All state and action entries except planar position are held fixed. The heads agree on the spatial value distribution, indicating low ensemble uncertainty for the critic correction at this rollout state.
Figure 8: Local critic analysis for the same slice as Figure 7 . The learned value remains consistent with the stage cost by avoiding low-value regions far from the goal, but shifts the local optimum according to the fixed velocity, attitude, and payload state. The critic is smooth in the plotted region, and the value-head variance defines the guidance scale used to blend the critic with the nominal MPC objective.
H
Nominal SR
Nominal cost
Critic SR
Critic cost
2
0.00
–
0.03
278
5
0.116
280
0.685
266
10
0.318
254
0.885
223
20
0.361
251
0.941
220
30
0.298
253
0.018
335
50
0.0
411
0.02
421
Appendix
Table 1: Influence of iLQR horizon length on cube-stacking performance at Δtctrl=0.05s . Nominal MPC requires an intermediate horizon, whereas the learned critic maintains strong performance even for short horizons.
Figure 9: Cube-stacking phases. (1) Approach and Grasp; (2) Pickup and Transport; (3) Align and Drop; Depth (B) and haptic observations (contact-probes in gripper fingers visualized by red markers, figure A1..3) improve grasp estimation. The predicted cube trajectory (C2, blue) matches the predicted end-effector trajectory (C2, green) during the transport phase (2).
Critic input
SR
Mean cost
[x,u,pgoal]
0.230
238.0
[x,u,pgoal,depth]
0.543
234.0
[x,u,pgoal,depth,haptic]
0.941
257.0
Appendix
Table 2: Influence of critic input features on cube-stacking performance. Adding depth and haptic observations improves success rate monotonically, indicating that richer observations improve grasp-phase inference and critic stability.
Experiment
SR
Mean cost
Baseline
0.394
234.7
g˙←[x,u,contact]
0.341
275.6
g˙←[contact,depth]
0.401
278.2
g˙←[x,u,contact,depth]
0.323
273.4
x˙←[x,u,contact]
0.361
296.4
Appendix
Table 3: Influence of residual dynamics learning on cube-stacking performance. Residual variants do not improve significantly over the nominal baseline and often increase cost.
Setting
Value
Planning time step Δtplan
0.05s
Horizon H
20
Parallel environments
100
iLQR iterations
8
Maximum line-search trials
9
Maximum regularization updates
10
Appendix
Table 4: Shared solver, optimization, and data-collection settings used in both tasks.
Setting
Quadrotor
Cube stacking
Episode length
5s
8s
Control time step Δtctrl
0.005s
0.01s
Residual model
ELU MLP, 3 layers, width 128
ELU MLP, 3 layers, width 256
Critic trunk
ELU MLP, 3 layers, width 128
ELU MLP, 3 layers, width 256
Value ensemble
4 heads, 1 layer, width 128 , mean aggregation
4 heads, 1 layer, width 256 , mean aggregation
Appendix
Table 5: Task-specific timing and network settings.
Autonomous Intelligent Systems, Computer Science Institute VI - Intelligent Systems and Robotics, Center for Robotics and the Lamarr Institute for Machine Learning and Artificial Intelligence, University of Bonn, Germany.