Despite rapid progress in vision-language-action (VLA) models, existing reasoning paradigms still face a fundamental \emph{state-representation mismatch} in open-loop planning. Given only an initial observation, models must internally simulate action-conditioned state transitions, whereas text-, pixel-, and latent-space reasoning can suffer from lossy spatial compression, error-accumulating visual generation, and bypass of intermediate latent tokens, respectively, undermining reliable long-horizon planning. We propose \textbf{State-Space Visual Reasoning} (SSVR), which decouples static visual context, language constraints, and a recurrent latent state. SSVR encodes the initial image and instruction once, then conditions each action prediction on the latent state and updates it with an action-conditioned GRU. Using Qwen2.5-VL as the backbone, SSVR achieves 99.5/99.6, 96.3/98.0, and 83.9/90.6 EM/PR on FrozenLake, Maze, and MiniBehavior, substantially outperforming prior methods. Extensive experiments support the effectiveness of recurrent state modeling for VLA open-loop planning across input transformations and transfer settings. By reusing static visual-textual context and updating a compact recurrent state, SSVR supports efficient multi-step inference, achieving up to 98.58× faster Maze decoding rollouts than the evaluated baselines with the prefix cache prebuilt.
Figures & tables
Figure 1: Motivation and evidence for aligning reasoning with dynamic state. Existing paradigms represent evolving state through compressed language, repeatedly generated images, or intermediate latent tokens. SSVR follows the finger-tracing intuition: a static scene is reused, a recurrent latent state is updated by actions through a GRU, and visual attention varies across planning steps.
Figure 2: Representative failure modes of existing reasoning paradigms. Text-space reasoning can hallucinate Maze topology after compressing the image into language; repeatedly generated pixel-space states may duplicate or erase objects and propagate these errors; latent-space reasoning can bypass both image and intermediate latent tokens when producing the final answer.
Figure 3: Overview of SSVR. The initial image and instruction are encoded once and kept as cached visual-textual context. Visual features initialize the latent state z0 . At step t , the model appends the state token et=zt , predicts action at , and updates the latent state through an GRU.
Model
VQA Support
FrozenLake
Maze
MiniBehavior
Avg.
Text Space
Gemini 2.0 Flash Direct ∗
✔
21.2 / 47.6
8.3 / 31.4
0.7 / 29.8
10.1 / 36.3
Gemini 2.0 Flash CoT ∗
✔
27.6 / 52.5
6.9 / 29.8
4.0 / 31.2
12.8 / 37.8
Gemini 2.5 Pro ∗
✔
72.0 / 85.0
21.5 / 35.5
37.6 / 59.9
43.7 / 60.1
Qwen2.5-VL-7B Direct ∗
✔
1.2 / 15.0
0.6 / 14.5
0.3 / 9.8
0.7 / 13.1
Qwen2.5-VL-7B CoT ∗
✔
8.2 / 29.1
2.3 / 15.2
0.5 / 14.7
3.7 / 19.7
Table 1: Comparison across reasoning paradigms. Entries report EM/PR (%). VQA Support: native VQA through text-based responses; ∗ : official release without training on our datasets.
Figure 4: Maze EM and PR across image scale ratios 0.7–1.3, overall and by difficulty level.
Table 6: Native Maze action-choice VQA in generation mode.
α
Maze EM
Maze PR
Entropy
VQAv2 Acc.
Paraphrase EM
Scale 0.7 EM
Scale 1.3 EM
1.0
91.8
95.7
0.0045
67.6
91.7
68.7
90.5
0.7
96.4
98.1
0.4833
68.0
96.8
79.4
96.4
0.4
95.5
97.5
0.7287
71.1
94.7
80.0
95.1
Table 7: Soft-target coefficient ablation on Maze.
Variant
Maze EM (%)
Maze PR (%)
VQAv2 (%)
SSVR
96.3
98.0
68.0
SSVR-head
96.4 +0.1
98.5 +0.5
46.9 -21.1
Table 8: Action-output ablation.
Figure 5: Gradient-weighted visual attribution from the GRU state-token query to image tokens. As the recurrent latent state changes, attribution highlights next-action-relevant regions: nearby walls in Maze, nearby holes in FrozenLake, and a phase-dependent shift from printer to table in MiniBehavior.
Method
Step latency (ms)
Rollout latency (ms)
Throughput (cases/s)
Peak memory (GiB)
SSVR †
45.611 (1.00x)
182.442 (1.00x)
5.481 (1.00x)
15.918 (1.00x)
SSVR
89.159 (1.95x)
356.634 (1.95x)
2.804 (0.51x)
15.937 (1.00x)
LVR
529.596 (11.61x)
1059.193 (5.81x)
0.944 (0.17x)
17.609 (1.11x)
Monet
1743.966 (38.24x)
6975.866 (38.24x)
0.143 (0.03x)
15.705 (0.99x)
VPRL
4496.111 (98.58x)
17984.442 (98.58x)
0.056 (0.01x)
13.331 (0.84x)
Table 9: Efficiency on Maze case0 . † : KV cache; latency excludes preprocessing and, for SSVR † , prefix-cache construction. Multipliers are relative to SSVR † .
Appendix figures & tables4 assets
Supplementary material from the paper’s appendix.
Appendix
Figure 6: Additional FrozenLake and Maze trajectories. In FrozenLake (left), attribution emphasizes nearby holes and safe local passages as the GRU latent state evolves with the predicted actions. In Maze (right), it moves across the fixed map and concentrates on relevant walls and openings. The shift across steps is consistent with latent-state-conditioned reading of a persistent scene.
Figure 7: Additional MiniBehavior trajectories. Each row follows a complete two-phase plan: the agent first navigates to the printer and executes PICK , then navigates to the table and executes DROP . The GRU updates its latent state from the predicted action history. Attribution correspondingly shifts from the printer and its approach path to the table and its approach path.
Task
Model state
Action tokens
FrozenLake
zt∈R3584
U/D/L/R
Maze
zt∈R3584
U/D/L/R
MiniBehavior
zt∈R3584
U/D/L/R/PICK/DROP
Appendix
Table 10: Dataset-specific state and action interfaces.
Setting
Value
Backbone
Qwen2.5-VL-7B-Instruct
Trainable adaptation
LoRA + GRU state module
LoRA rank / lora_alpha / dropout
32 / 64 / 0.1
LoRA target modules
q/k/v/o_proj , gate/down/up_proj
Optimizer
fused AdamW
Peak learning rate
1.5×10−4
Appendix
Table 11: Main training and evaluation configuration. The learning-rate scheduler is constructed for ten epochs, while training is stopped after epoch five.
Most Vision-Language-Action (VLA) models map observations directly to actions without explicit reasoning, limiting their capacity for reasoning-intensive long-horizon tasks. To address this, existing approaches adopt Chain-of-Thought (CoT) reasoning to enable subgoal decomposition and spatial anticipation. However, those methods lack a unified architecture for effective cross-modal reasoning and fail to explicitly include inverse reasoning ability based on the target state. We argue that manipulation planning naturally decomposes into prediction, anticipating the next visual state, and inverse dynamics, inferring the actions to reach it. Bridging both requires a unified autoregressive architecture that interleaves textual and visual reasoning in a single generation process. We propose \textbf{ThinkingVLA}, a generative model that realizes this decomposition within a unified Mixture-of-Transformers architecture. ThinkingVLA consists of a forward CoT that identifies the immediate subgoal and guides the visual forecasting; the predicted image then serves as the target state, grounding an inverse CoT that reasons about spatial relationships and action intent based on the predicted image; and the final action is generated conditioned on this full reasoning context. Extensive experiments on simulation and real-world benchmarks demonstrate that ThinkingVLA consistently outperforms state-of-the-art baselines, with particularly large gains on long-horizon manipulation tasks.
Tianyi Lu, Hui Zhang, Zijie Diao +8
Fudan University · Shanghai Innovation Institue · China Unicom +1
Recent work has begun to equip vision-language-action (VLA) policies with explicit intermediate reasoning. In embodied control, however, textual chain-of-thought is a poor fit: irrelevant or weakly textual information can interfere with action prediction, while autoregressive text decoding adds too much latency for real-time closed-loop execution. We present VISUALTHINK-VLA, a visual intermediate-reasoning framework for accurate, low-latency VLA policies. Our bootstrapping philosophy is to guide action with effective visual thinking: VISUALTHINK-VLA bootstraps action prediction through a compact visual-evidence interface that preserves spatial precision while avoiding decoding overhead. Besides, to further improve performance and efficiency, VISUALTHINK-VLA adopts a tailored selective routing mechanism to learn the visual evidence tokens, enabling low-latency inference while preserving high-capacity specialization. We also introduce VisualEvidence-Kit, a supervision-and-audit resource centered on a VisualEvidence-Agent that constructs a 754.7k VLA instructions VisualEvidence-Set for route supervision and counterfactual faithfulness tests. Across multiple benchmarks and real-robot evaluation, VISUALTHINK-VLA achieves the highest success rate on most benchmarks while reducing the multi-second latency of reasoning-augmented baselines to the sub-second regime. For example, on BridgeData V2, it reduces step latency from 8.377,s with ECoT to 0.367,s, achieving a 22.8 times speedup.
Mingjian Gao, Wenqiao Zhang, Yuqian Yuan +9
1Zhejiang University · 2Cornell University · 3National University of Singapore +1
While Vision-Language-Action (VLA) models have advanced embodied AI, their fundamentally reactive paradigm severely limits performance in partially observable and long-horizon tasks. When restricted to a single wrist-mounted camera, they inevitably suffer from perception forgetting as objects exit the field of view, and temporal task-progress forgetting} during multi-step execution. To overcome these bottlenecks, we propose AtlasVLA, a novel framework that transitions from direct reactive manipulation to proactive reasoning through a persistent world-ego state. AtlasVLA features a dual-memory architecture: a 4D Persistent World State Memory that lifts transient 2D observations into a globally updated, voxel-hashed spatial state to resolve visual blind spots, and an Ego-Working State Memory that tracks historical ego state and task progress. By conditioning a diffusion transformer (DiT) on this joint World-Ego state, AtlasVLA enables robust spatial reasoning. Extensive evaluations across LIBERO, RLBench, and real-world benchmarks demonstrate that AtlasVLA achieves state-of-the-art performance using solely a wrist camera. Remarkably, it decisively outperforms multi-view baselines, yielding absolute success rate improvements of 9.4% on LIBERO-Long and 17.5% in real-world long-horizon tasks.
Guiyu Zhao, Longteng Guo, Yanghong Mei +7
Institute of Automation, Chinese Academy of Sciences, Beijing, China · University of Chinese Academy of Sciences, Beijing, China · Beijing Freedo Technology Co., Ltd. +1