In CEM, a World Model Is Also a Proposal Mechanism
Organizations: UNSW Sydney · Harz University of Applied Sciences
Abstract
The cross-entropy method (CEM) uses world-model scores to select action sequences and fit the distribution sampled in its next iteration. A scoring error can therefore change both the present decision and the candidates considered later. We evaluate these two roles separately. Four types of predictive model generate CEM traces, and every model rescores every saved candidate pool. Executing the same candidates in the environment provides a reference elite set and proposal update. Across twelve independently trained task-seed units on Walker and Cheetah, the pre-specified proposal distance falls from the first to the final CEM iteration in every unit. Proposal widths contract and fitted means separate relative to the remaining search width. Pairwise ranking agreement stays near chance on Walker and declines on Cheetah; elite-set agreement does not improve. This comparison shows greater variation between scorers than between pool sources on Cheetah; Walker has variation in both and in their pairings. We use the original six units to select Random nonlinear for a one-update intervention, without inspecting intervention outcomes. Replacing its first model-ranked update with an environment-ranked update lowers final realised selected-sequence cost in those six units and in six further units held out from the selection.
Figures & tables
| Walker | Cheetah | |||
|---|---|---|---|---|
| Measurement | Iter. 0 | Iter. 3 | Iter. 0 | Iter. 3 |
| Mean distance | 0.200 | 0.288 | 0.184 | 0.292 |
| Log-width distance | 0.293 | 0.165 | 0.303 | 0.156 |
| Linear-width distance | 0.165 | 0.147 | 0.160 | 0.138 |
| Whitened distance | 0.260 | 0.324 | 0.244 | 0.324 |
| First-action proposal distance | 0.349 | 0.326 | 0.261 | 0.252 |
Appendix figures & tables13 assets
Supplementary material from the paper’s appendix.
Appendix
| Walker | Cheetah | |||
| Metric | Iter. 0 | Iter. 3 | Iter. 0 | Iter. 3 |
| Proposal distance | 0.254 | 0.235 | 0.252 | 0.235 |
| Proposal width | 0.350 | 0.159 | 0.350 | 0.131 |
| Whitened distance | 0.260 | 0.324 | 0.244 | 0.324 |
| First-action gap (action units) | 0.349 | 0.218 | 0.273 | 0.184 |
| First-action gap (proposal SD) | 0.998 | 1.321 | 0.779 | 1.258 |
| Omitted bin | Task | Seed 0 | Seed 1 | Seed 2 | Pass | |
|---|---|---|---|---|---|---|
| All states | Walker walk | -0.009 | -0.035 | -0.015 | 0 | no |
| All states | Cheetah run | -0.023 | -0.022 | -0.016 | 0 | no |
| Walker walk | -0.011 | -0.027 | -0.011 | 0 | no | |
| Cheetah run | -0.030 | -0.025 | -0.019 | 0 | no | |
| Walker walk | -0.000 | -0.040 | -0.020 | 0 | no | |
| Cheetah run | -0.020 | -0.017 | -0.025 | 0 | no |
| Omitted bin | Walker support | Cheetah support | Cross-task self interaction | Support | Interaction |
|---|---|---|---|---|---|
| All states | none | none | Identity; Random nonlinear | no | yes |
| none | none | Random nonlinear | no | yes | |
| none | none | Identity; Random nonlinear | no | yes | |
| none | none | Identity; Random nonlinear | no | yes | |
| Random nonlinear | none | Random nonlinear | no | yes |
| Cohort | Task | Seed | |||||
|---|---|---|---|---|---|---|---|
| Selection | Walker walk | 0 | |||||
| 1 | |||||||
| 2 | |||||||
| Cheetah run | 0 | ||||||
| 1 | |||||||
| 2 |
| Omitted bin | W3 | W4 | W5 | C3 | C4 | C5 | Rule |
|---|---|---|---|---|---|---|---|
| None | Pass | ||||||
| 1 | Pass | ||||||
| 3 | Pass | ||||||
| 5 | Pass | ||||||
| 7 | Pass |
| Audit pattern | Reading and next check |
|---|---|
| Scorer columns dominate | Test a score change or ranking loss against realised pairwise order and elite recovery. |
| Pool-source rows dominate | Test a proposal constraint or actor-anchored sampling. |
| Positive diagonal residual | Run a targeted model–planner intervention. |
| Metrics split as width contracts | Report width, ranking, and action-level fidelity. |
| Measurement | Walker | Cheetah |
|---|---|---|
| Initial mean displacement (action units) | 0.116 | 0.137 |
| Final mean displacement (action units) | 0.005 | 0.006 |
| Final width-floor occupancy (%) | 1.2 | 1.4 |
| Item | Setting |
|---|---|
| Runtime | Python 3.11.3; PyTorch 2.10.0; dm_control 1.0.37; MuJoCo 3.5.0; CUDA 12.8.0. |
| Dreamer architecture | Encoder channels 32–64–128–256; categorical RSSM with deterministic width 256 and stochastic categories; context width 1280; actor and critic hidden width 256. |
| Dreamer training | 500k environment steps per task–seed; five prefill episodes; replay capacity 500k; Adam with learning rates (world model) and (actor, critic), no weight decay; batches of eight length-16 sequences; one update per environment step; gradient clip 100. KL, reconstruction, reward, and continuation weights are 1; free nats is 1. The imagination horizon is 15, with , , and actor entropy weight . The final latest.pt checkpoint is frozen. |
| Surrogate data | 32 deterministic frozen-actor episodes per task–seed, shuffled with NumPy seed 0 and split 26/3/3 for training/validation/holdout; length-32 sequences; mean latent inference; rollout targets through horizon 15. |
| Lift networks | Identity: width 1280. Learned nonlinear: . Anchored nonlinear: affine concatenated with . Random nonlinear: fixed . MLP hidden layers use SiLU; outputs are linear. |
| Shared surrogate layers | Transition . Compressed-family decoder ; Identity uses the identity decoder. Reward and continuation heads are . |