cs.ROOct 8, 2026

PLaW-VLA: Predictive Latent World Modeling for Vision-Language-Action Policies

Authors: Yu Liu, Hetian Guo, Tianlv Huang, Ziyi Cai, Wudi Chen, Hantang Wang, Qiutong Liu, Yingzhi Peng, +6 more

Organizations: Jilin University · Astribot · Harbin Institute of Technology, Shenzhen · The Hong Kong Polytechnic University · The University of Tokyo

Abstract

Learning to predict how the world evolves can provide vision-language-action (VLA) policies with predictive context for long-horizon control, but its effectiveness depends on what future representation is modeled and how it conditions action generation. We introduce PLaW-VLA, which models task-relevant future states in a pretrained prediction-oriented representation space, reducing the need to predict control-irrelevant visual details. Built on a Mixture-of-Transformers architecture, PLaW-VLA conditions action generation on observation history, current task semantics, and predicted future states through structured causal attention. Experiments show a +11.8 percentage-point (pp) gain over reactive policies on RoboTwin Hard Horizon III and a +1.77 pp gain over reconstruction-oriented latent prediction on zero-shot LIBERO-Plus, supporting improved long-horizon control and generalization under distribution shift, respectively. By avoiding low-level visual reconstruction, PLaW-VLA lowers the burden of future prediction, enabling a lightweight latent world model with parallel future prediction and about 1/19 the inference latency of generative world-action modeling at comparable policy performance.

Explore similar work

CardsList
  1. World-Value-Action Model: Implicit Planning for Vision-Language-Action Systems

    Apr 16, 2026Runze Li, Hongyin Zhang, Junxi Jin +5Vision-Language-Action ModelsModel-Based Planning

  2. LaWAM: Latent World Action Models for Efficient Dynamics-Aware Robot Policies

    Jun 14, 2026Jialei Chen, Kai Wang, Kang Chen +9Latent Action LearningAction-Conditioned World Models

  3. Where Predictive Supervision Goes Shapes What VLA Policies Learn

    Sep 29, 2026Hanseul Kim, Jewon Yeom, Youngjoon Jeong +2Visuomotor Policy LearningVision-Language-Action Models