World-Model Policy Arbiter for Goal-Conditioned Reinforcement Learning
Organizations: University of Toronto · University of Toronto, Vector Institute
Abstract
Offline goal-conditioned reinforcement learning (GCRL) has produced a diverse set of goal-reaching algorithms, yet no single algorithm performs best across environments, goals, and even different phases of the same task. Rather than deploying only the best-performing policy, we ask whether a set of frozen goal-conditioned policies can be used collectively as a portfolio, deciding at every state which policy should act. Choosing a policy at each state is not straightforward. The policies' own value functions cannot be compared directly: they may use different scales, and some policies have no value function. We need to judge each policy by the states it is likely to reach, even though we can execute only one policy at a time. We also need to avoid switching so often that control becomes unstable. To address these challenges, we introduce World-Model Policy Arbiter (WMPA), a test-time framework that, given a set of frozen policies as input, rolls out each frozen policy in a learned state-space world model, evaluates the imagined futures with a shared goal-conditioned value function, and executes the highest-scoring policy for a short commitment interval before the next round of arbitration (policy selection). WMPA assumes access to a bank of frozen goal-conditioned policies and requires neither policy retraining nor privileged task-specific knowledge. Under the official OGBench evaluation protocol on 18 state-based datasets spanning maze navigation as well as cube, scene, and puzzle manipulation, WMPA improves the macro-average success rate from the 44% achieved by the best policy selected per dataset to 58%, with statistically significant gains on 12 datasets. These gains include +33 percentage points on cube-double-play and +36 percentage points on scene-play.
Figures & tables
| Family | Dataset | GCBC | GCIVL | GCIQL | QRL | CRL | HIQL | Best-Policy | WMPA |
|---|---|---|---|---|---|---|---|---|---|
| Maze | pointmaze-medium-navigate | ||||||||
| antmaze-large-navigate | |||||||||
| Cube | cube-single-play | ||||||||
| cube-single-noisy | |||||||||
| cube-double-play | |||||||||
| cube-double-noisy |
Appendix figures & tables13 assets
Supplementary material from the paper’s appendix.
Appendix
| Family | datasets | value head | WM | ||||
|---|---|---|---|---|---|---|---|
| Maze | 2 | metric | 1 | 10 | 0.999 | 0.9 | 10 |
| Cube | 6 | metric | 5 | 5 | 0.99 | 0.7 | 0 |
| Scene | 2 | metric | 10 | 5 | 0.998 | 0.7 | 10 |
| Puzzle | 8 | direct | 10 | 5 | 0.99 | 0.9 | – |
| success rate (%) | switches per episode | latency (ms per step) | |||||||||
|---|---|---|---|---|---|---|---|---|---|---|---|
| Dataset | Best-Policy | ||||||||||
| cube-double-play | |||||||||||
| scene-play | |||||||||||
| Dataset | Best | Random-Switch | WMPA | |
|---|---|---|---|---|
| pointmaze-medium-navigate | ||||
| antmaze-large-navigate | ||||
| cube-single-play | ||||
| cube-single-noisy | ||||
| cube-double-play | ||||
| cube-double-noisy |
| stall-restart, window | ||||||
|---|---|---|---|---|---|---|
| Dataset | Best | WMPA | ||||
| pointmaze-medium-navigate | ||||||
| antmaze-large-navigate | ||||||
| cube-single-play | ||||||
| cube-single-noisy | ||||||
| cube-double-play | ||||||
| Dataset | of WMPA | Best | every-step WM action search on Best | WMPA |
|---|---|---|---|---|
| pointmaze-medium-navigate | ||||
| antmaze-large-navigate | ||||
| cube-single-play | ||||
| cube-single-noisy | ||||
| cube-double-play | ||||
| cube-double-noisy |
| Dataset | Best | WMPA | true-sim. branches | (sim WMPA ) | |
|---|---|---|---|---|---|
| pointmaze-medium-navigate | +0 [-0, +1] | ||||
| antmaze-large-navigate | +4 [-0, +9] | ||||
| cube-single-play | +2 [-3, +7] | ||||
| cube-single-noisy | +0 [+0, +1] | ||||
| cube-double-play | +3 [-1, +7] | ||||
| cube-double-noisy | +0 [-10, +9] |
| WMPA (imagined states) | Q-select (no rollout) | ||||||
|---|---|---|---|---|---|---|---|
| Dataset | Best | metric value | direct value | (direct Q-select, ) | |||
| pointmaze-medium-navigate | † | +3 [+1, +4]* | |||||
| antmaze-large-navigate | † | +2 [-2, +6] | |||||
| cube-single-play | † | +18 [+2, +44]* | |||||
| cube-single-noisy | † | +0 [-1, +1] | |||||
| cube-double-play | † | -0 [-5, +5] | |||||
| Dataset | Best member | WMPA -own | WMPA -shared | |
|---|---|---|---|---|
| pointmaze-medium-navigate | ||||
| antmaze-large-navigate | ||||
| cube-single-play | ||||
| cube-single-noisy | ||||
| cube-double-play | ||||
| cube-double-noisy |
| Dataset | metric | direct | direct metric [95% CI] | head in Table 1 |
|---|---|---|---|---|
| cube-double-play | [ , ]* | metric | ||
| scene-play | [ , ]* | metric | ||
| puzzle-4x4-play | [ , ]* | direct | ||
| cube-triple-play | [ , ]* | metric |
| horizon agg. | ensemble agg. | ||||||
|---|---|---|---|---|---|---|---|
| Dataset | Best | WMPA | last | mean | min | mean-std | |
| pointmaze-medium-navigate | |||||||
| antmaze-large-navigate | |||||||
| cube-single-play | |||||||
| cube-single-noisy | |||||||
| cube-double-play | |||||||
| Dataset | Bank | Best member | WMPA | vs best member | vs six-member WMPA | |
|---|---|---|---|---|---|---|
| cube-double-play | Full bank | 6 | GCIQL ( ) | |||
| Pooled play+noisy | 12 | GCIQL ( ) | +43 [+35, +50]* | +10 [+6, +15]* | ||
| scene-play | Full bank | 6 | GCIQL ( ) | |||
| Pooled play+noisy | 12 | GCIQL ( ) | +37 [+30, +43]* | +1 [-3, +5] | ||
| puzzle-4x4-play | Full bank | 6 | GCIQL ( ) | |||
| Pooled play+noisy | 12 | GCIQL (noisy) ( ) | +18 [+11, +26]* | +2 [-4, +8] |
| Dataset | (oracle) | WMPA | WMPA [95% CI] | WMPA [95% CI] | |||
|---|---|---|---|---|---|---|---|
| pointmaze-medium-navigate | +3 [-17, +17] | -19 [-22, -16]* | |||||
| antmaze-large-navigate | +0 [-3, +4] | -9 [-12, -6]* | |||||
| cube-single-play | +14 [+10, +18]* | +4 [-1, +9] | |||||
| cube-single-noisy | +1 [+0, +2]* | +0 [-0, +1] | |||||
| cube-double-play | +33 [+25, +40]* | +12 [+6, +18]* | |||||
| cube-double-noisy | +35 [+30, +40]* | +29 [+24, +34]* |
| switches/ep | ms/step (CPU) | |||||||||
|---|---|---|---|---|---|---|---|---|---|---|
| Dataset | ep. len | WMPA | Rand. | decisions/ep | WM calls/step | usage (bits) | WMPA | Rand. | fixed | |
| pointmaze-medium-navigate | 295 | 24.8 | 246.5 | 294.8 | 18.0 | 1.30 | 3.76 | 0.21 | 0.18 | |
| antmaze-large-navigate | 475 | 335.9 | 425.8 | 475.2 | 18.0 | 2.57 | 3.88 | 0.19 | 0.15 | |
| cube-single-play | 90 | 13.0 | 27.3 | 18.3 | 18.3 | 2.54 | 3.29 | 0.35 | 0.23 | |
| cube-single-noisy | 46 | 3.7 | 13.0 | 9.5 | 18.8 | 1.61 | 3.24 | 0.24 | 0.30 | |
| cube-double-play | 264 | 37.8 | 73.8 | 53.1 | 18.1 | 2.55 | 2.97 | 0.23 | 0.22 | |