Organizations: Department of Electrical and Computer Engineering, National University of Singapore, Singapore 119077, Singapore · Department of Mechanical Engineering, National University of Singapore, Singapore 119077, Singapore
Continuous legged manipulation requires accurate end-effector tracking while the base keeps walking. Combining reinforcement learning (RL) with model predictive control (MPC) suits this task: the learned policy provides robust locomotion, while MPC coordinates the base and arm to compensate for tracking errors. However, MPC can compensate only for base motion that it can predict, and a learned policy's command response varies with gait phase, contact, and payload. We present ReCo, a framework that couples response-consistent locomotion with policy-aware MPC for legged manipulation. Response shaping trains the policy to respond to commands consistently and repeatably across randomized dynamics. An identified closed-loop response model then lets MPC jointly plan locomotion commands and arm motion. On the simulation benchmark, ReCo reduces position and orientation root-mean-square error (RMSE) by 28.7% and 27.4% relative to the best baseline for each metric. Real-world experiments demonstrate onboard continuous legged manipulation with coordinated base and arm motion.
Figures & tables
Fig. 1: Demonstration of continuous legged manipulation tasks in the real world. ReCo jointly plans locomotion commands and arm motions to track task-space end-effector trajectories for object organization, cabinet-door opening, and pick-and-place.
Fig. 2: Overview of ReCo. Response shaping and model identification produce a closed-loop response model of the locomotion policy. MPC integrates this model, arm kinematics, and state feedback to jointly plan locomotion commands and arm motion.
Term
Definition and purpose
Reference tracking
rref in ( 4 ); matches the commanded transient.
Phase repeatability
ℓphase=∑iμiphwiph(yi−yˉi−δ^i)2 ; regularizes the gait waveform.
ℓdomaind in ( 5 ); matches randomized instances to the nominal response.
TABLE I: Response-shaping terms. The reference reward contributes to r+ ; the three nonnegative losses enter r− with negative weights.
Controller
Paradigm
Position RMSE (m) ↓
Orientation RMSE ( ∘ ) ↓
Progress (%) ↑
Success ↑
Falls ↓
Ma et al. [ 9 ]
RL+MPC
0.199 ± 0.114
32.7 ± 14.6
82.6
31/100
8/100
Pure MPC [ 25 ]
Model-based
0.189 ± 0.038
46.8 ± 14.2
80.5
19/100
12/100
UMI on Legs [ 16 ]
Learning
0.883 ± 0.832
81.1 ± 34.3
12.9
12/100
80/100
RoboDuet [ 8 ] +PP
Learning
0.679 ± 0.603
79.9 ± 29.0
65.7
42/100
9/100
Visual WBC [ 15 ] +PP
Learning
1.454 ± 1.101
68.2 ± 30.1
25.0
23/100
38/100
ReCo (Ours)
RL+MPC
0.135 ± 0.091
23.8 ± 13.7
94.0
75/100
3/100
TABLE II: Uneven-terrain MuJoCo benchmark: 50 trajectories under nominal and push conditions (100 runs per controller). Errors are per-run RMSE means ± sample standard deviations, including recorded prefixes of early failures. PP denotes the external omnidirectional pure-pursuit tracker.
Fig. 3: Representative benchmark trajectories. Top: 3D EE paths colored by normalized time. Bottom: reference phase rate s˙ and path curvature κ (clipped at the 95th percentile for display). Spatial extent, curvature, and traversal timing vary across trajectories.
Fig. 4: Selected controllers on one nominal benchmark trajectory. (a) Final recorded states, which occur at different times after early termination; yellow and cyan mark reference and actual EE paths. (b) Equal-scale XY paths and tracking errors (orientation in radians); curves use a five-sample moving average; RMSE uses unfiltered samples. UMI on Legs terminates after a fall.
Response shaping
–
–
✓
✓
Model identification
–
✓
–
✓
Position RMSE (m)
0.257
0.232
0.166
0.135
Orientation RMSE ( ∘ )
21.7
24.3
21.5
23.8
EE jerk (m/s 3 )
252.1
250.1
261.0
249.7
Base jerk (m/s 3 )
145.6
140.3
155.4
149.9
Leg torque RMS (N ⋅ m)
6.41
6.40
6.46
6.30
TABLE III: Factorial ablation over response shaping and model identification. Without identification, MPC assumes instantaneous command execution. All metrics are lower-is-better; the best entry per row is bold.
Condition
Unshaped
Shaped
Normalized prediction MSE
Horizon 0.25 s
0.0958
0.0395
Horizon 0.50 s
0.1091
0.0442
Horizon 1.00 s
0.1097
0.0434
Forward-velocity residual RMS (m/s)
Arm held
0.0135
0.0099
TABLE IV: Prediction accuracy and response repeatability. Response shaping lowers prediction error at every horizon and reduces the sensitivity of the forward-velocity response to arm motion.
Fig. 5: Forward-velocity prediction for the unshaped and shaped policies, labeled standard and response-consistent in the legend. Measured responses and fixed 0.50 s forecasts are plotted at forecast target times on identical axes. Table IV reports aggregate errors.
Fig. 6: Phase-conditioned forward-velocity residual in a selected moving-arm run. Table IV reports mean residual RMS across the matched conditions.
Fig. 7: Gait-periodic height prediction. (a) Measured height relative to the nominal offset and fixed 0.25 s forecasts. (b) Height RMSE across 36 trajectories, including a shifted-phase control. The reported aggregate uses 0.10, 0.25, and 0.34 s; 0.50 and 1.00 s are also shown. All forecasts use causal phase extrapolation.
Fig. 8: Representative trajectories from the motion-capture experiments. Left: hardware execution. Right: reference and measured EE paths.
In humanoid motion control, model predictive control (MPC) offers physically grounded prediction and constraint handling, while reinforcement learning (RL) enables robust whole-body skills through large-scale simulation. However, using MPC inside RL often requires time-consuming problem construction or excessive training overhead, making such frameworks difficult to justify in practice. This work studies efficient training-time MPC guidance for humanoid locomotion and manipulation, termed MPC-RL. We introduce a centroidal-dynamics MPC reward formulation that leverages guidance from MPC trajectories in training time. To make this practical in massively parallel RL, we develop πnMPC, a parallel-in-horizon and construction-free batched GPU MPC solver that operates directly on time-varying dynamics to avoid high memory usage and pre-compilation. Through a variety of comparative studies and hardware validations, we have found that MPC-RL achieves superior performance in locomotion and manipulation skills. The code base is available at https://github.com/junhengl/mpc-rl.
Junheng Li, Liang Wu, Sergio A. Esteban +3
California Institute of Technology, CA 91106, USA · Johns Hopkins University, MD 21218, USA
Real-world reinforcement learning (RL) offers a promising route to dexterous manipulation policies that can adapt directly from physical interaction, but learning is hindered by inefficient early exploration and costly failures. We propose a framework that uses sampling-based model predictive control (MPC) as scaffolding for real-world dexterous RL, providing structured prior experience and task-directed guidance during learning without human demonstrations or corrective actions. A small set of MPC trajectories is first used to populate an offline replay buffer and to pretrain the actor and critic. During online learning, MPC intermittently guides data collection while an off-policy Soft Actor-Critic learner trains from both prior MPC experience and newly collected physical interaction, with control gradually transitioning to the learned policy. On continuous in-hand rotation with a 16-DoF Allegro hand, initialized from 20 MPC trajectories collected in 12 minutes on hardware, the policy reaches 100% success after 7 minutes of online RL, with about three object drops on average during training. After 20 minutes of online learning, the policy achieves more than five times the rotation speed of the MPC controller and completes 1000 consecutive rotations without a drop. Ablations show complementary benefits from MPC-based pretraining, retained MPC experience, and online MPC guidance, while additional experiments demonstrate rapid adaptation to new object geometries and successful goal-conditioned reorientation.
Emek Barış Küçüktabak, Karankumar Patel, Zhaodong Yang +3
Honda Research Institute USA, San Jose, CA, USA · Skild AI, San Mateo, CA, USA · Georgia Institute of Technology, Atlanta, GA, USA
Bipedal loco-manipulation enables robots to interact with objects beyond the nominal workspace of their arms by coordinating locomotion and manipulation. Realizing this capability requires a low-level whole-body controller that translates task-level manipulation goals into coordinated arm and leg motions while maintaining balance. We present a unified whole-body controller trained with reinforcement learning that directly maps 6-DoF end-effector targets to coordinated actions for the bipedal base and robotic arm. Given only an end-effector target, the learned controller autonomously coordinates reaching, postural adaptation, and stepping without explicit base-velocity or footstep commands. A reward-gating strategy regulates the trade-offs among end-effector tracking, locomotion, and balance during training, while a temporal context estimator combines windowed Transformer encoding, recurrent GRU memory, and auxiliary dynamics prediction to extract dynamics-relevant information from observation history. Real-robot experiments demonstrate that the same controller supports reaching, postural adaptation, and stepping under commands from VR teleoperation, a learned diffusion policy, and scripted trajectories, providing a common end-effector interface for diverse manipulation tasks.