cs.CVSep 27, 2026

Resolving State-Representation Mismatch: State-Space Visual Reasoning for Open-Loop VLA Planning

Authors: Junhao Xiao, Haoxiang Zhao, Menghao Fang, Jinkui Zhang, Jinghan Yu, Xinyu Huang, Zhiyu Wu, Kaiming Xu, +3 more

Organizations: FDU · HUST · TJU · CCNU · Kuaishou

Abstract

Despite rapid progress in vision-language-action (VLA) models, existing reasoning paradigms still face a fundamental \emph{state-representation mismatch} in open-loop planning. Given only an initial observation, models must internally simulate action-conditioned state transitions, whereas text-, pixel-, and latent-space reasoning can suffer from lossy spatial compression, error-accumulating visual generation, and bypass of intermediate latent tokens, respectively, undermining reliable long-horizon planning. We propose \textbf{State-Space Visual Reasoning} (SSVR), which decouples static visual context, language constraints, and a recurrent latent state. SSVR encodes the initial image and instruction once, then conditions each action prediction on the latent state and updates it with an action-conditioned GRU. Using Qwen2.5-VL as the backbone, SSVR achieves 99.5/99.6, 96.3/98.0, and 83.9/90.6 EM/PR on FrozenLake, Maze, and MiniBehavior, substantially outperforming prior methods. Extensive experiments support the effectiveness of recurrent state modeling for VLA open-loop planning across input transformations and transfer settings. By reusing static visual-textual context and updating a compact recurrent state, SSVR supports efficient multi-step inference, achieving up to 98.58×98.58\times faster Maze decoding rollouts than the evaluated baselines with the prefix cache prebuilt.

Figures & tables

Appendix figures & tables4 assets

Supplementary material from the paper’s appendix.

Appendix

Explore similar work

CardsList
  1. ThinkingVLA: Interleaved Vision and Language Reasoning for Robotic Manipulation

    Jun 16, 2026Tianyi Lu, Hui Zhang, Zijie Diao +8Diffusion-Based Vision-Language-ActionsRobotic Manipulation

  2. VisualThink-VLA: Visual Intermediate Reasoning for Effective and Low-Latency Vision-Language-Action Policies

    May 28, 2026Mingjian Gao, Wenqiao Zhang, Yuqian Yuan +9Latent Visual ReasoningVisual Reasoning

  3. AtlasVLA: Persistent World-Ego State Modeling for Vision-Language-Action Models

    Aug 7, 2026Guiyu Zhao, Longteng Guo, Yanghong Mei +7Diffusion-Based Vision-Language-ActionsSpatial Reasoning