cs.AIOct 7, 2026

World Potential Model: Pretrained World Knowledge as Progress Potentials

Authors: Jun Zhao, Jixin Tang, Yang Shu, Jinyang Wu, Yuyang Lu, Jingqi Tong, Hao Xu, Weifeng Ge, +1 more

Organizations: National University of Singapore · Fudan University · Tsinghua University · The University of Sydney

Abstract

Long-horizon language agents often receive supervision only from terminal task outcomes, leaving little signal for distinguishing productive intermediate behavior from stagnation or even regression. Rather than learning a separate value function or process reward model for every task, we ask whether pretrained models can recognize task progress from their existing world knowledge. We formalize this capability with a World Potential Model (WPM), a goal-conditioned evaluator of task-relative realized progress in agent contexts. In ALFWorld and ScienceWorld, off-the-shelf pretrained models substantially outperform chance at recovering realized-progress structure without task-specific evaluator fine-tuning. We further anchor these progress judgments to task-specific milestones to obtain scalar world potentials, whose temporal differences provide process-sensitive step-level credit for policy optimization. Under matched comparisons, WPM-guided optimization improves success over outcome-only GRPO across all evaluated configurations. Together, these results provide initial evidence that pretrained world knowledge can support reusable realized-progress evaluation and provide useful supervision for long-horizon agents.

Figures & tables

Explore similar work

CardsList
  1. Self-Evolving World Models for LLM Agent Planning

    Jun 29, 2026Xuan Zhang, Wenxuan Zhang, See-Kiong Ng +1Test-Time AdaptationWorld Model Learning

  2. Qwen-AgentWorld: Language World Models for General Agents

    Jun 23, 2026Yuxin Zuo, Zikai Xiao, Li Sheng +30LLM World ModelsAgentic RL

  3. MemWM: Memory-Augmented Text-Based World Model

    Aug 7, 2026Yujun Wang, Tao Zhang, Jinhe Bi +9LLM World ModelsWorld Model Learning