World Potential Model: Pretrained World Knowledge as Progress Potentials
Organizations: National University of Singapore · Fudan University · Tsinghua University · The University of Sydney
Abstract
Long-horizon language agents often receive supervision only from terminal task outcomes, leaving little signal for distinguishing productive intermediate behavior from stagnation or even regression. Rather than learning a separate value function or process reward model for every task, we ask whether pretrained models can recognize task progress from their existing world knowledge. We formalize this capability with a World Potential Model (WPM), a goal-conditioned evaluator of task-relative realized progress in agent contexts. In ALFWorld and ScienceWorld, off-the-shelf pretrained models substantially outperform chance at recovering realized-progress structure without task-specific evaluator fine-tuning. We further anchor these progress judgments to task-specific milestones to obtain scalar world potentials, whose temporal differences provide process-sensitive step-level credit for policy optimization. Under matched comparisons, WPM-guided optimization improves success over outcome-only GRPO across all evaluated configurations. Together, these results provide initial evidence that pretrained world knowledge can support reusable realized-progress evaluation and provide useful supervision for long-horizon agents.
Figures & tables
| Model | ALFWorld | ScienceWorld | Avg. | ||
|---|---|---|---|---|---|
| Pairwise | Anchored | Pairwise | Anchored | ||
| Random Guess | 50.00 | 20.00 | 50.00 | 15.27 | 33.82 |
| MiMo-VL-7B XiaomiMiMo (2025) | 92.67 | 58.00 | 97.33 | 63.33 | 77.83 |
| Qwen2.5-VL-7B Qwen Team (2025b) | 80.00 | 60.00 | 81.33 | 48.00 | 67.33 |
| Qwen2.5-32B Qwen Team (2024) | 95.33 | 80.67 | 96.67 | 76.00 | 87.17 |
| Qwen2.5-VL-72B Qwen Team (2025a) | 93.33 | 82.00 | 94.67 | 81.33 | 87.83 |
| Model | Method | Alf World | Science World |
|---|---|---|---|
| Qwen2.5‑3B | Vanilla | 10.93 | 8.59 |
| GRPO | 53.90 | 51.56 | |
| WPM | 60.50 | 55.39 | |
| Qwen2.5‑7B | Vanilla | 15.62 | 13.28 |
| GRPO | 68.70 | 66.40 | |
| WPM | 72.10 | 74.73 |