cs.CVSep 28, 2026

WorldGuide: Learning Success-Failure Boundaries in Latent World Models for Vision-Language-Action Policies

Authors: Lin Liu, Lu Zhang, Ziying Song, Wu Yang, Yuzheng Zhuang, Yunzhi Zhuge, Shuai Tao, Wulong Liu, +1 more

Organizations: School of Information and Communication Engineering Dalian University of Technology & Beta Infinity · School of Information and Communication, Engineering Dalian University of Technology · Nanyang Technological University · Beta Infinity

Abstract

Latent world models offer a promising way to improve Vision-Language-Action policies by capturing the consequences of actions. However, models trained primarily on expert demonstrations have limited exposure to failure outcomes and may struggle to distinguish visually similar successful and failed interactions. We propose \textbf{WorldGuide}, a framework that learns these distinctions in latent space and uses them to guide policy training. WorldGuide combines predictive pretraining on successful and failed trajectories with contrastive learning on matched success--failure pairs. The learned predictor then provides a differentiable reward to guide joint optimization of the policy and visual encoder. The predictor is discarded after training, so deployment requires no additional world-model inference. Extensive experiments show that WorldGuide substantially improves VLA reliability and achieves state of the art performance on LIBERO 100 and SimplerEnv, reaching \textbf{96.8%} and \textbf{72.0%}, respectively. Code will be publicly available.

Figures & tables

Appendix figures & tables9 assets

Supplementary material from the paper’s appendix.

Appendix

Explore similar work

CardsList
  1. World Pilot: Steering Vision-Language-Action Models with World-Action Priors

    Jun 10, 2026Zefu Lin, Rongxu Cui, Junjia Xu +4Efficient World-Action ModelWorld Models

  2. WoVR: World Models as Reliable Simulators for Post-Training VLA Policies with RL

    Feb 15, 2026Zhennan Jiang, Shangqing Zhou, Yutong Jiang +11Video World ModelsSimulation-Based Reinforcement Learning

  3. Think Like a World Model, Act Like a VLA: Distilling World-Model Representations into Compact Robot Policies

    Sep 21, 2026Trung Dao, Sankalp Yamsani, Jaden Park +2World ModelsRobot Systems