Planning with learned world models combines online trajectory optimization with learned value and policy functions for high-dimensional control. Because the planner determines the experience used for learning, while the learned critic and actor in turn score and propose future plans, planning and learning form a closed feedback loop. TD-MPC is a prominent instance of this design. Recent policy-constrained variants strengthen one part of the loop by aligning the learned policy with planner behavior. We introduce PL-MPC (Planning-Learning MPC), which additionally modifies critic supervision and planner terminal-value estimation. Hybrid multi-step TD targets expose critic updates to more realized rewards before bootstrapping; disagreement-aware terminal estimates reduce the influence of uncertain critic values during MPPI planning; and return-weighted actor distillation emphasizes planner-executed actions from high-return episodes. The world-model architecture and MPPI optimizer are otherwise unchanged. On HumanoidBench, the largest gains occur on \texttt{balance-hard}, where Total Average Return (TAR) increases from 98±18 to 387±255, and \texttt{hurdle}, from 199±13 to 466±200; performance across the broader benchmark remains task dependent, and PL-MPC remains competitive on DMControl. Controlled ablations show different component interactions across the two tasks. We further demonstrate zero-shot sim-to-real transfer on wrench-nut alignment with a 7-DoF KUKA IIWA14, obtaining higher observed success than TD-M(PC)^2 on the training object size and two unseen sizes. Code and data will be available at: https://pl-mpc-humanoid.github.io.
Figures & tables
Fig. 1: Zero-shot real-robot wrench–nut alignment. (a) PL-MPC is trained entirely in Isaac Gym and deployed on a KUKA IIWA14 without real-world fine-tuning; the deployed checkpoint is selected within the first 500K training steps. (b) Representative trials on the training size (size 5) and unseen sizes 3 and 1. Frames are sampled at different times during each execution and are not consecutive control steps. In the shown size 5 and size 3 trials, PL-MPC completes the task while TD-M(PC) 2 reaches the 128-step limit. On size 1, both methods succeed, with PL-MPC completing the trial in fewer control steps. Successful sequences end with the wrench seated around the nut.
Fig. 2: PL-MPC architecture. The upper loop comprises MPPI planning (A), replay (B), world-model and value learning (C), and actor learning (D). Matching numbered badges link each modification to its detailed panel. 1: MTD constructs critic targets in C from replay rewards, using actor-driven model tails beyond the available replay slice; detail 1A shows the observed/model reward counts. 2: ATE penalizes the planner’s online random-average terminal value in A using target-critic ensemble disagreement. 3: RAD uses realized episode returns to weight distillation of planner-executed replay actions in D. The actor supplies planning proposals and actions for model tails and TD bootstraps. Arrows show information flow; locks denote detached targets or weights. Dotted and solid target-separation curves illustrate one-step and multi-step TD supervision, respectively; these curves, ensemble spreads, and value/weight diagrams are schematic.
Task
DreamerV3
TD-MPC2
BMPC
BOOM
TD-M(PC) 2
PL-MPC (Ours)
Target
DMControl
dog-stand
41±9
523±421
932±21
982±7
832±99
976±14
–
dog-trot
11±3
394±184
953±10
923±8
899±39
885±78
–
humanoid-stand
5±1
637±71
955±8
920±15
929±15
954±2
–
humanoid-walk
2±0
643±96
933±20
921±10
867±63
947±6
–
HumanoidBench Locomotion
Fig. 3: Benchmark performance. Total Average Return (TAR; mean ± std across seeds). Green and gray cells denote the highest and second-highest row means, respectively; ties share the best highlight. Target denotes the HumanoidBench reference-return threshold and is distinct from the environment-defined success metric.
balance-hard
hurdle
Variant
TAR
Succ.
TAR
Succ.
TD-M(PC) 2 3M
98±18
0.00
199±13
0.00
w/o MTD
186±51
0.00
530±275
0.47
w/o ATE
206±63
0.00
296±21
0.00
w/o RAD
272±221
0.15
231±72
0.00
PL-MPC (Ours)
387±255
0.19
466±200
0.19
Fig. 4: Leave-one-out component ablation. Total Average Return (TAR) and environment-defined success rate (Succ.), averaged over three seeds. Each variant removes one component from full PL-MPC; when enabled, RAD uses the default 200K warmup.
Fig. 5: Leave-one-out ablations and training dynamics. Each panel compares full PL-MPC with variants that remove MTD , ATE , or RAD . Top: TAR, environment-defined success, and per-seed return averaged over the final five evaluations; dotted lines denote HumanoidBench reference-return thresholds. Bottom: mean critic value, TD-target standard deviation, and RAD weight. Curves show mean ±1 standard deviation across seeds. The TD-M(PC) 2 final return is shown for reference.
Fig. 6: Targeted ablations on the two analysis tasks. (a) On balance-hard , ATE and RAD are fixed and only MTD is removed. (b) On hurdle , both variants use MTD ; the comparison adds ATE and RAD jointly. Top: episode return, task success, and per-seed return averaged over the final five evaluations; dotted lines denote HumanoidBench reference-return thresholds. Bottom: mean critic value, standard deviation of the TD target, and RAD weight. The TD-M(PC) 2 backbone is shown in the performance panels for reference. Diagnostic panels compare the corresponding PL-MPC variants. Curves show mean ±1 standard deviation across seeds.
Task
TD-M(PC) 2
PL-MPC (Ours)
200K
750K
Δ
Succ. 200K/750K
crawl
899±64
855±31
971±6
+116
0.99/0.99
maze
347±9
353±3
349±8
−4
0.00/0.00
walk
922±8
911±7
910±1
−1
0.99/0.99
balance-simple
599±389
765±176
809±32
+44
0.63/0.65
hurdle
199±13
466±200
414±127
−52
0.19/0.09
Fig. 7: Sensitivity to RAD warmup. Total Average Return (TAR; mean ± std across three seeds) for the TD-M(PC) 2 backbone and PL-MPC with 200K and 750K RAD warmups. Δ denotes the change from 200K to 750K; Succ. reports the environment-defined success rates for the two warmups.
Method
Size
N
SR (%)
95% CI
Steps
Pos. (mm)
Rot. (deg)
Training object size
TD-M(PC) 2
5
31
61.3
[43.8,76.3]
53.3±24.4
11.08±4.31
7.28±4.26
PL-MPC (Ours)
5
31
74.2
[56.8,86.3]
55.9±25.9
9.86±4.50
6.14±2.35
Unseen object sizes
TD-M(PC) 2
3
15
73.3
[48.0,89.1]
54.1±16.0
10.60±4.54
3.57±1.92
PL-MPC (Ours)
3
15
80.0
[54.8,93.0]
47.9±14.6
11.57±4.41
4.64±1.55
Fig. 8: Zero-shot real-robot wrench–nut alignment. Success rate (SR) is computed over all trials with 95% Wilson confidence intervals. Steps and final position/rotation errors are mean ± sample standard deviation over successful trials. Size 5 is used during simulation training; sizes 3 and 1 are unseen. Pooled rows combine the two unseen sizes. Bold green denotes the higher observed SR.
Humans solve complex problems by constructing plans and mentally simulating their outcomes with an internal model of the world. Machine learning has made substantial progress in learning world models that predict the consequences of action sequences, yet the procedures used to plan with these models remain largely hand-designed. Most planners rely on fixed search or optimization rules; approaches that learn aspects of search typically imitate a predefined optimizer or use planning to inform an amortized policy, rather improving multi-step plans. We introduce \textbf{Reinforced Planning}, a method that learns the plan-update itself by reinforcing update rules that produce better plans, using gradients propagated through a differentiable world model. We instantiate Reinforced Planning in RP1, which learns a critic over imagined outcomes via temporal-difference learning and a neural plan-improvement operator trained via imagined rollouts with a pretrained world model. RP1 can be trained fully offline without environment interaction; environment episodes are used only for checkpoint selection. Across visual navigation, arm reaching, and robotic manipulation on two world-model backbones, RP1 matches or exceeds existing planners, achieving near-perfect success in several settings while using 1,000× fewer world-model rollouts than the strongest alternative (CEM) and planning up to 67× faster under concurrent planners inference.
Planning with a learned latent world model is a promising route to control from raw pixels, but a strong world model alone is not enough. We show this experimentally: even with a perfect world model (operationalized by replacing the learned forward predictor with an idealized rollout of the true environment dynamics), a finite-budget sample-based planner still fails on some tasks, indicating that the bottleneck can lie in search rather than in world-model accuracy. Motivated by this gap, we propose IMWM (Intuition Model + World Model), which pairs the world model with an intuition model trained from demonstrations to recognize promising actions. The two models collaborate through three lightweight components: (i) Retrieval Initialization, which initializes the planner's action proposal from a retrieved demonstration; (ii) Hybrid Cost, which combines the intuition score with the world-model rollout cost; and (iii) a Reliability Gate, which adjusts how much the planner trusts intuition in each setting. Across four pixel-based goal-reaching tasks (Two-Room, Reacher, Push-T, and OGBench-Cube), IMWM has higher mean success than the world-model-only planner on all four, with the largest gains on Two-Room (99.2%, +11.5 percentage points) and OGBench-Cube (94.7%, +28.5 percentage points).
Baoqi Gao, Ruize Han, Miao Wang +1
Beihang University · Shenzhen University of Advanced Technology
Action-conditioned JEPA world models enable planning toward visually specified goals without reconstructing future pixels, yet latent prediction alone does not explicitly encourage the learned representations to retain information relevant to robotic control. We introduce an end-to-end JEPA world model that augments latent prediction with inverse dynamics (IDM) and state alignment (SA). While inverse dynamics discourages latent collapse and makes latent transitions informative of the actions that produced them, state alignment grounds consecutive representations in their associated physical configuration and motion. Across four benchmark tasks, our model attains the highest success rates on TwoRoom (100%), PushT (98%), and OGBench-Cube (87%), while performing comparably to LeWorldModel on Reacher. Our ablation further shows that adding state alignment consistently improves planning success over IDM alone across all four tasks. Although LeWorldModel, our primary baseline, attains higher average straightening on OGBench-Cube, transition-subspace analysis shows that its transition energy is concentrated in a substantially lower-dimensional subspace. Our state-aligned model exhibits a higher effective transition dimension than LeWorldModel and improves planning over IDM alone, supporting state alignment as an effective complement to inverse dynamics for robotic planning.