cs.AISep 30, 2026

Beyond Prediction: Steering VLM Agents with Retrospective World Modeling

Authors: Yongjiang Liu, Jie Zhang, Haoyue Zhang, Jingcai Guo, Deze Zeng, Song Guo

Organizations: The Hong Kong University of Science and Technology · The Hong Kong Polytechnic University · China University of Geoscience

Abstract

Equipping VLM agents with world modeling capabilities has shown strong potential for complex reasoning and long-horizon planning, while reducing the dependence of policy learning on costly real-world interactions. Existing methods mainly rely on prospective simulation to predict the consequences of candidate actions. However, this forward-only paradigm focuses on what will happen next and provides limited constraints for verifying whether an action is causally consistent with the observed state transition, which can lead to plausible-looking but physically incoherent behaviors. In this paper, we challenge the view of world modeling as only prospective prediction and introduce Retrospective World Modeling, a new agent learning paradigm that enables agents to reason backward by estimating the retrospective attribution distribution P(a^t∣st,st+1)P(\hat{a}{t}|s_t, s{t+1}) for the action that most likely caused a given transition. Based on this capability, we formulate the Self-Consistency Reward (SCR), an intrinsic signal that measures the probabilistic consistency between the policy action and the retrospective explanation. Integrating SCR into reinforcement learning provides dense transition-level feedback and steers agents toward behaviors that are both task-effective and physically grounded. Extensive experiments across diverse agentic tasks show that our method substantially improves policy robustness and generalization over prospective-only world modeling baselines.

Explore similar work

CardsList
  1. Self-Evolving World Models for LLM Agent Planning

    Jun 29, 2026Xuan Zhang, Wenxuan Zhang, See-Kiong Ng +1Large Language Model PlanningWorld Models

  2. Learning Implicit Causal World Models from Multi-Agent Demonstrations

    Jul 28, 2026Jasorsi GhoshEfficient World-Action ModelModel-Based Reinforcement Learning