Despite rapid progress in vision-language-action (VLA) models, existing reasoning paradigms still face a fundamental \emph{state-representation mismatch} in open-loop planning. Given only an initial observation, models must internally simulate action-conditioned state transitions, whereas text-, pixel-, and latent-space reasoning can suffer from lossy spatial compression, error-accumulating visual generation, and bypass of intermediate latent tokens, respectively, undermining reliable long-horizon planning. We propose \textbf{State-Space Visual Reasoning} (SSVR), which decouples static visual context, language constraints, and a recurrent latent state. SSVR encodes the initial image and instruction once, then conditions each action prediction on the latent state and updates it with an action-conditioned GRU. Using Qwen2.5-VL as the backbone, SSVR achieves 99.5/99.6, 96.3/98.0, and 83.9/90.6 EM/PR on FrozenLake, Maze, and MiniBehavior, substantially outperforming prior methods. Extensive experiments support the effectiveness of recurrent state modeling for VLA open-loop planning across input transformations and transfer settings. By reusing static visual-textual context and updating a compact recurrent state, SSVR supports efficient multi-step inference, achieving up to 98.58× faster Maze decoding rollouts than the evaluated baselines with the prefix cache prebuilt.
Figures & tables
Figure 1: Motivation and evidence for aligning reasoning with dynamic state. Existing paradigms represent evolving state through compressed language, repeatedly generated images, or intermediate latent tokens. SSVR follows the finger-tracing intuition: a static scene is reused, a recurrent latent state is updated by actions through a GRU, and visual attention varies across planning steps.
Figure 2: Representative failure modes of existing reasoning paradigms. Text-space reasoning can hallucinate Maze topology after compressing the image into language; repeatedly generated pixel-space states may duplicate or erase objects and propagate these errors; latent-space reasoning can bypass both image and intermediate latent tokens when producing the final answer.
Figure 3: Overview of SSVR. The initial image and instruction are encoded once and kept as cached visual-textual context. Visual features initialize the latent state z0 . At step t , the model appends the state token et=zt , predicts action at , and updates the latent state through an GRU.
Model
VQA Support
FrozenLake
Maze
MiniBehavior
Avg.
Text Space
Gemini 2.0 Flash Direct ∗
✔
21.2 / 47.6
8.3 / 31.4
0.7 / 29.8
10.1 / 36.3
Gemini 2.0 Flash CoT ∗
✔
27.6 / 52.5
6.9 / 29.8
4.0 / 31.2
12.8 / 37.8
Gemini 2.5 Pro ∗
✔
72.0 / 85.0
21.5 / 35.5
37.6 / 59.9
43.7 / 60.1
Qwen2.5-VL-7B Direct ∗
✔
1.2 / 15.0
0.6 / 14.5
0.3 / 9.8
0.7 / 13.1
Qwen2.5-VL-7B CoT ∗
✔
8.2 / 29.1
2.3 / 15.2
0.5 / 14.7
3.7 / 19.7
Table 1: Comparison across reasoning paradigms. Entries report EM/PR (%). VQA Support: native VQA through text-based responses; ∗ : official release without training on our datasets.
Figure 4: Maze EM and PR across image scale ratios 0.7–1.3, overall and by difficulty level.
Table 6: Native Maze action-choice VQA in generation mode.
α
Maze EM
Maze PR
Entropy
VQAv2 Acc.
Paraphrase EM
Scale 0.7 EM
Scale 1.3 EM
1.0
91.8
95.7
0.0045
67.6
91.7
68.7
90.5
0.7
96.4
98.1
0.4833
68.0
96.8
79.4
96.4
0.4
95.5
97.5
0.7287
71.1
94.7
80.0
95.1
Table 7: Soft-target coefficient ablation on Maze.
Variant
Maze EM (%)
Maze PR (%)
VQAv2 (%)
SSVR
96.3
98.0
68.0
SSVR-head
96.4 +0.1
98.5 +0.5
46.9 -21.1
Table 8: Action-output ablation.
Figure 5: Gradient-weighted visual attribution from the GRU state-token query to image tokens. As the recurrent latent state changes, attribution highlights next-action-relevant regions: nearby walls in Maze, nearby holes in FrozenLake, and a phase-dependent shift from printer to table in MiniBehavior.
Method
Step latency (ms)
Rollout latency (ms)
Throughput (cases/s)
Peak memory (GiB)
SSVR †
45.611 (1.00x)
182.442 (1.00x)
5.481 (1.00x)
15.918 (1.00x)
SSVR
89.159 (1.95x)
356.634 (1.95x)
2.804 (0.51x)
15.937 (1.00x)
LVR
529.596 (11.61x)
1059.193 (5.81x)
0.944 (0.17x)
17.609 (1.11x)
Monet
1743.966 (38.24x)
6975.866 (38.24x)
0.143 (0.03x)
15.705 (0.99x)
VPRL
4496.111 (98.58x)
17984.442 (98.58x)
0.056 (0.01x)
13.331 (0.84x)
Table 9: Efficiency on Maze case0 . † : KV cache; latency excludes preprocessing and, for SSVR † , prefix-cache construction. Multipliers are relative to SSVR † .
Appendix figures & tables4 assets
Supplementary material from the paper’s appendix.
Appendix
Figure 6: Additional FrozenLake and Maze trajectories. In FrozenLake (left), attribution emphasizes nearby holes and safe local passages as the GRU latent state evolves with the predicted actions. In Maze (right), it moves across the fixed map and concentrates on relevant walls and openings. The shift across steps is consistent with latent-state-conditioned reading of a persistent scene.
Figure 7: Additional MiniBehavior trajectories. Each row follows a complete two-phase plan: the agent first navigates to the printer and executes PICK , then navigates to the table and executes DROP . The GRU updates its latent state from the predicted action history. Attribution correspondingly shifts from the printer and its approach path to the table and its approach path.
Task
Model state
Action tokens
FrozenLake
zt∈R3584
U/D/L/R
Maze
zt∈R3584
U/D/L/R
MiniBehavior
zt∈R3584
U/D/L/R/PICK/DROP
Appendix
Table 10: Dataset-specific state and action interfaces.
Setting
Value
Backbone
Qwen2.5-VL-7B-Instruct
Trainable adaptation
LoRA + GRU state module
LoRA rank / lora_alpha / dropout
32 / 64 / 0.1
LoRA target modules
q/k/v/o_proj , gate/down/up_proj
Optimizer
fused AdamW
Peak learning rate
1.5×10−4
Appendix
Table 11: Main training and evaluation configuration. The learning-rate scheduler is constructed for ten epochs, while training is stopped after epoch five.
Institute of Automation, Chinese Academy of Sciences, Beijing, China · University of Chinese Academy of Sciences, Beijing, China · Beijing Freedo Technology Co., Ltd. +1