Visual world models enable goal-directed control through decision-time action search, but their deployment efficiency is often limited by conservatively large planning budgets. We show that competitive task performance can be achieved without agreement with the Full-budget action, that sufficient budgets vary across model--task pairs, and that iterative planners repeatedly encode solve-invariant context. To address these inefficiencies, we propose {SufficientPlan}, a simple deployment framework that requires no modification to pretrained world models or planner updates. Its {Paired Sequential Budget Certification (PSBC)} component uses paired closed-loop evidence to search for and certify a reduced model--task-specific budget within a predefined Full-performance tolerance. Its {Static-Context Reuse (SCR)} component caches observation and goal representations across search iterations while preserving candidate-dependent planning and selected actions. Experiments across multiple world-model backbones and visual-control tasks show that SufficientPlan substantially reduces search budgets and planning latency while maintaining competitive control performance.
Figures & tables
Figure 1: Task success without Full-budget action agreement. On INTACT Push-T, the calibrated and Full planners follow different task- and action-space trajectories but both achieve 96% success, using 2,400 and 9,000 scored sequences per solve, respectively.
Figure 2: Paired Sequential Budget Certification (PSBC). PSBC compares the current budget with the full-budget reference on paired calibration episodes. Uncertain cases receive additional paired evidence, insufficient budgets are increased, the first proposal satisfying the certification criterion is returned as B⋆ and frozen for held-out deployment.
Figure 3: Static-Context Reuse (SCR). SCR encodes the solve-invariant observation history and goal once and reuses them across K CEM iterations, reducing cost-stage static-encoder calls from 2K to 2 without changing candidate-dependent planning or CEM updates.
Figure 4: Matched-budget evaluation of Static-Context Reuse (SCR). (a) Matched-budget success rates and episode-level outcome consistency. (b) Mean planning-latency reduction. (c) Planning speedup across backbone–task settings. SCR preserves all evaluated outcomes while achieving up to 2.18× speedup.
Planner
Deployment
Push-T
Cube
Reacher
TwoRoom
Budget
Success
Budget
Success
Budget
Success
Budget
Success
LeWM
Full CEM
Full
85
Full
71
Full
82
Full
91
LeWM+PSBC
Calibrated CEM
50.0%
88
13.3%
74
80.0%
82
30.0%
91
Fast-LeWM
Full CEM
Full
95
Full
68
Full
82
Full
94
Fast-LeWM+PSBC
Calibrated CEM
26.7%
89
13.3%
73
80.0%
81
50.0%
94
PRISM
Official PoG-MPPI
Full
88
Full
78
–
–
–
–
Table 1: Comparison of full-budget planning and PSBC-calibrated planning across different world-model planners and control tasks. Budget denotes the planning budget relative to the corresponding full planner, and Success denotes the task success rate (%). “Full” indicates the original full planning budget, while “–” denotes unavailable results.
Figure 5: SCR scaling with planning depth on Fast-LeWM/Push-T. (a) Mean planning latency of the uncached and SCR-enabled planners. (b) Measured speedup and analytical reduction in solve-static encoder calls. Repeated measurements show no reliable speedup at one iteration, whereas SCR reaches 2.45× speedup at 30 iterations as repeated static encoding becomes more prominent.
Figure 6: Paired qualitative comparison on Cube. From the same initial state and goal, LeWM+PSBC succeeds while Full LeWM fails; Ground Truth is shown as reference, and columns progress from left to right.
Backbone
Configuration
PSBC
SCR
Scored Seq./Solve
Budget (% Full)
SR (%)
Δ SR (pp)
INTACT
Full Actor-CEM
✗
✗
9,000
100.0
96
0
+ PSBC
✓
✗
2,400
26.7
96
0
+ PSBC + SCR
✓
✓
2,400
26.7
96
0
Fast-LeWM
Full CEM
✗
✗
9,000
100.0
90
0
+ PSBC
✓
✗
2,400
26.7
89
−1
+ PSBC + SCR
✓
✓
2,400
26.7
89
−1
Table 2: Incremental ablation on the frozen Push-T development set. PSBC reduces the deployed search budget while retaining near-Full success, and SCR preserves the resulting budget and control outcome.
Appendix figures & tables13 assets
Supplementary material from the paper’s appendix.
Appendix
Figure A: Global fixed budgets versus PSBC-selected budgets. (a) Numbers of model–task pairs for which each global budget is below, equal to, or above the PSBC-selected budget. (b) Selected budget fractions across the 16 evaluated pairs.
Quantity
Value
Full planning budget
9,000
PSBC deployment budget
600
Full time per episode
34.28 s
PSBC time per episode
10.93 s
Time saving per episode
23.35 s
Appendix
Table A: Deployment efficiency on PRISM/Cube. PSBC reduces per-episode planning time relative to Full planning.
Figure B: Success–budget profiles across world-model planners and tasks. Lines show development-set success rates across planning budgets, while diamonds and stars denote held-out Full and PSBC results, respectively. Development curves and held-out markers correspond to distinct evaluation splits.
Task
Seed
B⋆
Stop n
qB,n
qB,100
SR: B /Full
Push-T
0
3,300
20
0.956
0.510
88/90
Push-T
1
1,800
20
0.956
0.799
88/87
Cube
0
900
80
0.972
0.971
88/82
Cube
1
600
100
0.981
0.981
92/85
Appendix
Table B: Posterior evolution after PSBC certification on PRISM. q100 retrospectively evaluates the locked 100-episode set and is not used in the original stopping decision.
Budget
% Full
SR
Full SR
Δ SR
Rescued
Broken
qB,100
Empirical Accept
Fixed- n Cert.
300
3.3
83
85
−2
10
12
0.503
✓
✗
600
6.7
92
85
+7
13
6
0.981
✓
✓
1,200
13.3
89
85
+4
12
8
0.911
✓
✗
2,400
26.7
91
85
+6
10
4
0.984
✓
✓
Appendix
Table C: Budget-selection criteria on PRISM/Cube. All candidates use the same 100 paired development episodes with a 2-point tolerance; fixed- n certification requires qB,100≥0.95 .
Figure C: Task success and action refinement on INTACT Push-T. (a) Development success rate across 100 episodes at different CEM iterations. (b) Mean RMS distance between the first action block at each iteration and its 30-iteration counterpart, with paired bootstrap 95% confidence intervals. Eight and 30 iterations achieve the same 96% success rate despite selecting different actions.
Deployment
Actor Prior
PSBC
SCR
Seq./Solve
Budget (% Full)
Resulting SR (%)
Δ SR vs Direct
Direct
✓
✗
✗
0
0.0
83
0
Guarded-A
✓
✗
✗
384
4.3
88
+5
Pure CEM
✗
✗
✗
9,000
100.0
92
+9
Full Actor-CEM
✓
✗
✗
9,000
100.0
96
+13
Actor-CEM + PSBC
✓
✓
✗
2,400
26.7
96
+13
SufficientPlan
✓
✓
✓
2,400
26.7
96
+13
Appendix
Table D: INTACT deployment ablation on the frozen Push-T development set. We isolate the effects of actor-guided initialization, PSBC, and SCR.
Figure D: Sequential certification trace of PSBC on PRISM/Cube. Curves show the posterior probability that each candidate budget is non-inferior to Full planning within the prescribed tolerance as paired calibration episodes accumulate. At 100 episodes, the 600-sequence proposal crosses the 0.95 certification threshold and is returned as the certified operating budget.
Figure E: Multi-seed success rates on Push-T and Cube. Bars and error bars report the mean and standard deviation over evaluation seeds {0,1,42} , and markers show the individual results. Percentages above the PSBC bars denote the corresponding deployment budgets relative to Full planning.
Figure F: Paired qualitative results on Reacher, TwoRoom, and Push-T. From identical initial states and goals, LeWM+PSBC succeeds while Full LeWM fails; rows correspond to the three tasks and columns progress from left to right.
τ
δ (pp)
0.90
0.95
0.99
0
100
–
–
2
100
100
–
5
80
80
100
Appendix
Table E: Sensitivity of the 600-sequence proposal on PRISM/Cube. Entries report the paired-episode inspection at which the proposal is first certified. “–” indicates that the proposal is not certified within 100 paired episodes.
Budget
Iter.
Uncached (s)
SCR (s)
Static Enc. Reduction
Latency Reduction
Speedup
300
1
0.4521
0.4571
0.0%
−1.1 %
0.99 ×
600
2
1.2422
0.8000
50.0%
35.6%
1.55 ×
1,200
4
2.0664
1.0660
75.0%
48.4%
1.94 ×
2,400
8
3.4617
1.7080
87.5%
50.7%
2.03 ×
4,500
15
6.2457
2.6330
93.3%
57.8%
2.37 ×
9,000
30
11.4736
4.6796
96.7%
59.2%
2.45 ×
Appendix
Table F: SCR scaling on Fast-LeWM/Push-T. SCR provides no reliable speedup at one iteration, while its latency benefit increases with search depth.
Contract
INTACT
Fast-LeWM
Candidate population/hash equal
✓
✓
Candidate cost equal
✓
✓
CEM mean/std updates equal
✓
✓
Final selected action equal
✓
✓
Solver RNG state equal
✓
✓
Model buffers unchanged
✓
✓
Appendix
Table G: Full-budget numerical equivalence contract for SCR. Uncached and SCR-enabled planners are evaluated with common random numbers at the maximum 9,000-sequence budget. The contract verifies the internal planning trajectory rather than only aggregate task success.
Planning with a learned latent world model is a promising route to control from raw pixels, but a strong world model alone is not enough. We show this experimentally: even with a perfect world model (operationalized by replacing the learned forward predictor with an idealized rollout of the true environment dynamics), a finite-budget sample-based planner still fails on some tasks, indicating that the bottleneck can lie in search rather than in world-model accuracy. Motivated by this gap, we propose IMWM (Intuition Model + World Model), which pairs the world model with an intuition model trained from demonstrations to recognize promising actions. The two models collaborate through three lightweight components: (i) Retrieval Initialization, which initializes the planner's action proposal from a retrieved demonstration; (ii) Hybrid Cost, which combines the intuition score with the world-model rollout cost; and (iii) a Reliability Gate, which adjusts how much the planner trusts intuition in each setting. Across four pixel-based goal-reaching tasks (Two-Room, Reacher, Push-T, and OGBench-Cube), IMWM has higher mean success than the world-model-only planner on all four, with the largest gains on Two-Room (99.2%, +11.5 percentage points) and OGBench-Cube (94.7%, +28.5 percentage points).
Baoqi Gao, Ruize Han, Miao Wang +1
Beihang University · Shenzhen University of Advanced Technology
Token-based world models enable fine-grained latent planning, but repeatedly processing large spatial token grids makes action search expensive. We introduce COSTGRAD, a training-free, goal-conditioned selector that ranks spatial tokens by the gradient norm of the planning cost with respect to each input token. By deriving importance from the downstream control objective, COSTGRAD targets tokens that matter for planning rather than merely for prediction. On AdaLN-conditioned predictors at 50% sparsity, COSTGRAD matches or exceeds full-token planning on three of four continuous-control benchmarks, while giving a measured 2.6× wall-clock speedup per environment planning step. Combining token sparsity with reduced CEM search increases this to a ∼5× total speedup while still exceeding the full-token baseline. We also identify an architecture-dependent failure mode: in a matched AdaLN-vs-concat comparison, concat maintains comparable full-token performance but pure COSTGRAD loses its advantage over random selection. This difference tracks action-pathway drift: gradient-selected removal produces less drift than random removal on AdaLN, but more on concat. These results highlight selector-architecture compatibility as a design axis for sparse world-model planning. Project page and demos: https://ycxuyingchen.github.io/costgrad/
Different research lines use the term world model in different ways, yet they share a common aim: to capture how the world evolves under action in a form that supports perception, simulation, and planning. Two prominent realizations are neural predictors that learn dynamics in continuous vector spaces, and hand-built physics engines that expose explicit state and physical laws. Neural predictors scale from data but leave the form of the dynamics implicit; physics engines are inspectable and editable but difficult to construct at scale. We introduce VisualPatchWorld (VPW), which represents world dynamics as code. VPW first selects a qualitative dynamical form with short active probes, then fits that form's free parameters from recorded state-action traces by minimizing multi-step prediction error. The resulting programs can be rolled forward like a simulator, inspected in source form, and used inside model-predictive control; image-derived scene graphs can supply the live state at replan time. Across comparisons with prior code-based world models, VPW attains 69.0% mean planning success and exceeds the strongest code baseline by 23.5 points. The largest gains arise when choosing the correct qualitative dynamics is essential. Under the same planner, the induced models approach ground-truth engine success on navigation and grasp-rich control; a residual gap remains for contact-rich pushing, and checking a shortlist of promising plans in the engine closes most of that gap. These results establish a practical route toward automatically constructed code world models that are useful for planning. Code is available at https://github.com/HKBU-KnowComp/VisualPatchWorld/.