Organizations: Department of Electrical and Computer Engineering, National University of Singapore, Singapore 119077, Singapore · Department of Mechanical Engineering, National University of Singapore, Singapore 119077, Singapore
Continuous legged manipulation requires accurate end-effector tracking while the base keeps walking. Combining reinforcement learning (RL) with model predictive control (MPC) suits this task: the learned policy provides robust locomotion, while MPC coordinates the base and arm to compensate for tracking errors. However, MPC can compensate only for base motion that it can predict, and a learned policy's command response varies with gait phase, contact, and payload. We present ReCo, a framework that couples response-consistent locomotion with policy-aware MPC for legged manipulation. Response shaping trains the policy to respond to commands consistently and repeatably across randomized dynamics. An identified closed-loop response model then lets MPC jointly plan locomotion commands and arm motion. On the simulation benchmark, ReCo reduces position and orientation root-mean-square error (RMSE) by 28.7% and 27.4% relative to the best baseline for each metric. Real-world experiments demonstrate onboard continuous legged manipulation with coordinated base and arm motion.
Figures & tables
Fig. 1: Demonstration of continuous legged manipulation tasks in the real world. ReCo jointly plans locomotion commands and arm motions to track task-space end-effector trajectories for object organization, cabinet-door opening, and pick-and-place.
Fig. 2: Overview of ReCo. Response shaping and model identification produce a closed-loop response model of the locomotion policy. MPC integrates this model, arm kinematics, and state feedback to jointly plan locomotion commands and arm motion.
Term
Definition and purpose
Reference tracking
rref in ( 4 ); matches the commanded transient.
Phase repeatability
ℓphase=∑iμiphwiph(yi−yˉi−δ^i)2 ; regularizes the gait waveform.
ℓdomaind in ( 5 ); matches randomized instances to the nominal response.
TABLE I: Response-shaping terms. The reference reward contributes to r+ ; the three nonnegative losses enter r− with negative weights.
Controller
Paradigm
Position RMSE (m) ↓
Orientation RMSE ( ∘ ) ↓
Progress (%) ↑
Success ↑
Falls ↓
Ma et al. [ 9 ]
RL+MPC
0.199 ± 0.114
32.7 ± 14.6
82.6
31/100
8/100
Pure MPC [ 25 ]
Model-based
0.189 ± 0.038
46.8 ± 14.2
80.5
19/100
12/100
UMI on Legs [ 16 ]
Learning
0.883 ± 0.832
81.1 ± 34.3
12.9
12/100
80/100
RoboDuet [ 8 ] +PP
Learning
0.679 ± 0.603
79.9 ± 29.0
65.7
42/100
9/100
Visual WBC [ 15 ] +PP
Learning
1.454 ± 1.101
68.2 ± 30.1
25.0
23/100
38/100
ReCo (Ours)
RL+MPC
0.135 ± 0.091
23.8 ± 13.7
94.0
75/100
3/100
TABLE II: Uneven-terrain MuJoCo benchmark: 50 trajectories under nominal and push conditions (100 runs per controller). Errors are per-run RMSE means ± sample standard deviations, including recorded prefixes of early failures. PP denotes the external omnidirectional pure-pursuit tracker.
Fig. 3: Representative benchmark trajectories. Top: 3D EE paths colored by normalized time. Bottom: reference phase rate s˙ and path curvature κ (clipped at the 95th percentile for display). Spatial extent, curvature, and traversal timing vary across trajectories.
Fig. 4: Selected controllers on one nominal benchmark trajectory. (a) Final recorded states, which occur at different times after early termination; yellow and cyan mark reference and actual EE paths. (b) Equal-scale XY paths and tracking errors (orientation in radians); curves use a five-sample moving average; RMSE uses unfiltered samples. UMI on Legs terminates after a fall.
Response shaping
–
–
✓
✓
Model identification
–
✓
–
✓
Position RMSE (m)
0.257
0.232
0.166
0.135
Orientation RMSE ( ∘ )
21.7
24.3
21.5
23.8
EE jerk (m/s 3 )
252.1
250.1
261.0
249.7
Base jerk (m/s 3 )
145.6
140.3
155.4
149.9
Leg torque RMS (N ⋅ m)
6.41
6.40
6.46
6.30
TABLE III: Factorial ablation over response shaping and model identification. Without identification, MPC assumes instantaneous command execution. All metrics are lower-is-better; the best entry per row is bold.
Condition
Unshaped
Shaped
Normalized prediction MSE
Horizon 0.25 s
0.0958
0.0395
Horizon 0.50 s
0.1091
0.0442
Horizon 1.00 s
0.1097
0.0434
Forward-velocity residual RMS (m/s)
Arm held
0.0135
0.0099
TABLE IV: Prediction accuracy and response repeatability. Response shaping lowers prediction error at every horizon and reduces the sensitivity of the forward-velocity response to arm motion.
Fig. 5: Forward-velocity prediction for the unshaped and shaped policies, labeled standard and response-consistent in the legend. Measured responses and fixed 0.50 s forecasts are plotted at forecast target times on identical axes. Table IV reports aggregate errors.
Fig. 6: Phase-conditioned forward-velocity residual in a selected moving-arm run. Table IV reports mean residual RMS across the matched conditions.
Fig. 7: Gait-periodic height prediction. (a) Measured height relative to the nominal offset and fixed 0.25 s forecasts. (b) Height RMSE across 36 trajectories, including a shifted-phase control. The reported aggregate uses 0.10, 0.25, and 0.34 s; 0.50 and 1.00 s are also shown. All forecasts use causal phase extrapolation.
Fig. 8: Representative trajectories from the motion-capture experiments. Left: hardware execution. Right: reference and measured EE paths.