Planning with learned world models combines online trajectory optimization with learned value and policy functions for high-dimensional control. Because the planner determines the experience used for learning, while the learned critic and actor in turn score and propose future plans, planning and learning form a closed feedback loop. TD-MPC is a prominent instance of this design. Recent policy-constrained variants strengthen one part of the loop by aligning the learned policy with planner behavior. We introduce PL-MPC (Planning-Learning MPC), which additionally modifies critic supervision and planner terminal-value estimation. Hybrid multi-step TD targets expose critic updates to more realized rewards before bootstrapping; disagreement-aware terminal estimates reduce the influence of uncertain critic values during MPPI planning; and return-weighted actor distillation emphasizes planner-executed actions from high-return episodes. The world-model architecture and MPPI optimizer are otherwise unchanged. On HumanoidBench, the largest gains occur on \texttt{balance-hard}, where Total Average Return (TAR) increases from 98±18 to 387±255, and \texttt{hurdle}, from 199±13 to 466±200; performance across the broader benchmark remains task dependent, and PL-MPC remains competitive on DMControl. Controlled ablations show different component interactions across the two tasks. We further demonstrate zero-shot sim-to-real transfer on wrench-nut alignment with a 7-DoF KUKA IIWA14, obtaining higher observed success than TD-M(PC)^2 on the training object size and two unseen sizes. Code and data will be available at: https://pl-mpc-humanoid.github.io.
Figures & tables
Fig. 1: Zero-shot real-robot wrench–nut alignment. (a) PL-MPC is trained entirely in Isaac Gym and deployed on a KUKA IIWA14 without real-world fine-tuning; the deployed checkpoint is selected within the first 500K training steps. (b) Representative trials on the training size (size 5) and unseen sizes 3 and 1. Frames are sampled at different times during each execution and are not consecutive control steps. In the shown size 5 and size 3 trials, PL-MPC completes the task while TD-M(PC) 2 reaches the 128-step limit. On size 1, both methods succeed, with PL-MPC completing the trial in fewer control steps. Successful sequences end with the wrench seated around the nut.
Fig. 2: PL-MPC architecture. The upper loop comprises MPPI planning (A), replay (B), world-model and value learning (C), and actor learning (D). Matching numbered badges link each modification to its detailed panel. 1: MTD constructs critic targets in C from replay rewards, using actor-driven model tails beyond the available replay slice; detail 1A shows the observed/model reward counts. 2: ATE penalizes the planner’s online random-average terminal value in A using target-critic ensemble disagreement. 3: RAD uses realized episode returns to weight distillation of planner-executed replay actions in D. The actor supplies planning proposals and actions for model tails and TD bootstraps. Arrows show information flow; locks denote detached targets or weights. Dotted and solid target-separation curves illustrate one-step and multi-step TD supervision, respectively; these curves, ensemble spreads, and value/weight diagrams are schematic.
Task
DreamerV3
TD-MPC2
BMPC
BOOM
TD-M(PC) 2
PL-MPC (Ours)
Target
DMControl
dog-stand
41±9
523±421
932±21
982±7
832±99
976±14
–
dog-trot
11±3
394±184
953±10
923±8
899±39
885±78
–
humanoid-stand
5±1
637±71
955±8
920±15
929±15
954±2
–
humanoid-walk
2±0
643±96
933±20
921±10
867±63
947±6
–
HumanoidBench Locomotion
Fig. 3: Benchmark performance. Total Average Return (TAR; mean ± std across seeds). Green and gray cells denote the highest and second-highest row means, respectively; ties share the best highlight. Target denotes the HumanoidBench reference-return threshold and is distinct from the environment-defined success metric.
balance-hard
hurdle
Variant
TAR
Succ.
TAR
Succ.
TD-M(PC) 2 3M
98±18
0.00
199±13
0.00
w/o MTD
186±51
0.00
530±275
0.47
w/o ATE
206±63
0.00
296±21
0.00
w/o RAD
272±221
0.15
231±72
0.00
PL-MPC (Ours)
387±255
0.19
466±200
0.19
Fig. 4: Leave-one-out component ablation. Total Average Return (TAR) and environment-defined success rate (Succ.), averaged over three seeds. Each variant removes one component from full PL-MPC; when enabled, RAD uses the default 200K warmup.
Fig. 5: Leave-one-out ablations and training dynamics. Each panel compares full PL-MPC with variants that remove MTD , ATE , or RAD . Top: TAR, environment-defined success, and per-seed return averaged over the final five evaluations; dotted lines denote HumanoidBench reference-return thresholds. Bottom: mean critic value, TD-target standard deviation, and RAD weight. Curves show mean ±1 standard deviation across seeds. The TD-M(PC) 2 final return is shown for reference.
Fig. 6: Targeted ablations on the two analysis tasks. (a) On balance-hard , ATE and RAD are fixed and only MTD is removed. (b) On hurdle , both variants use MTD ; the comparison adds ATE and RAD jointly. Top: episode return, task success, and per-seed return averaged over the final five evaluations; dotted lines denote HumanoidBench reference-return thresholds. Bottom: mean critic value, standard deviation of the TD target, and RAD weight. The TD-M(PC) 2 backbone is shown in the performance panels for reference. Diagnostic panels compare the corresponding PL-MPC variants. Curves show mean ±1 standard deviation across seeds.
Task
TD-M(PC) 2
PL-MPC (Ours)
200K
750K
Δ
Succ. 200K/750K
crawl
899±64
855±31
971±6
+116
0.99/0.99
maze
347±9
353±3
349±8
−4
0.00/0.00
walk
922±8
911±7
910±1
−1
0.99/0.99
balance-simple
599±389
765±176
809±32
+44
0.63/0.65
hurdle
199±13
466±200
414±127
−52
0.19/0.09
Fig. 7: Sensitivity to RAD warmup. Total Average Return (TAR; mean ± std across three seeds) for the TD-M(PC) 2 backbone and PL-MPC with 200K and 750K RAD warmups. Δ denotes the change from 200K to 750K; Succ. reports the environment-defined success rates for the two warmups.
Method
Size
N
SR (%)
95% CI
Steps
Pos. (mm)
Rot. (deg)
Training object size
TD-M(PC) 2
5
31
61.3
[43.8,76.3]
53.3±24.4
11.08±4.31
7.28±4.26
PL-MPC (Ours)
5
31
74.2
[56.8,86.3]
55.9±25.9
9.86±4.50
6.14±2.35
Unseen object sizes
TD-M(PC) 2
3
15
73.3
[48.0,89.1]
54.1±16.0
10.60±4.54
3.57±1.92
PL-MPC (Ours)
3
15
80.0
[54.8,93.0]
47.9±14.6
11.57±4.41
4.64±1.55
Fig. 8: Zero-shot real-robot wrench–nut alignment. Success rate (SR) is computed over all trials with 95% Wilson confidence intervals. Steps and final position/rotation errors are mean ± sample standard deviation over successful trials. Size 5 is used during simulation training; sizes 3 and 1 are unseen. Pooled rows combine the two unseen sizes. Bold green denotes the higher observed SR.