Vision-language-action (VLA) models increasingly incorporate intermediate reasoning to improve robotic manipulation, yet existing approaches primarily reason about observed states without explicitly anticipating future scene evolution. Extending such reasoning to explicit future rollouts at every inference step, however, introduces substantial computational overhead. We propose IG-VLA, a VLA reasoning framework that enables models to imagine the future and internalize the gist. Our Latent Spatiotemporal Reasoning learns to imagine task-relevant future scene evolution directly in visual representation space, guiding action prediction without costly pixel-level video generation. To further reduce inference overhead, we introduce Scene Gist Memory, which internalizes reasoning-derived scene-behavior associations into a compact Scene Gist Token, preserving the benefits of future reasoning while bypassing explicit future imagination at inference. Extensive experiments on LIBERO, LIBERO-Plus, and VLABench demonstrate the effectiveness and efficiency of IG-VLA. On the LIBERO-Plus Language suite, both the reasoning and gist policies outperform the strongest baseline by nearly 6% in success rate. The gist policy also achieves up to 6.38x speedup over baselines, reducing inference latency from 1081ms to 169.5ms per action chunk on a single NVIDIA A6000 GPU. These results demonstrate that future spatiotemporal reasoning can be effectively internalized for efficient VLA deployment.
Figures & tables
Figure 1: Language-based vs. latent visual reasoning. (a) Language-based reasoning may lose fine-grained spatial information, leading to imprecise actions; (b) Latent visual reasoning preserves evolving scene geometry through visual future reasoning, enabling more precise actions.
Figure 2: Human Automaticity vs. Scene Gist Memory.
Figure 3: Overview of our framework. (a) the overall architecture; (b) the latent spatiotemporal reasoning module learns to predict future visual keyframes; (c) the gist modules internalize reasoning-derived behavioral information into a compact Scene Gist Token.
Methods
Spatial
Object
Goal
Long
Avg.
SR ↑
Rank ↓
SR ↑
Rank ↓
SR ↑
Rank ↓
SR ↑
Rank ↓
SR ↑
Rank ↓
Diffusion Policy
78.3
29
92.5
21
68.3
30
50.5
30
72.4
30
Octo
78.9
28
85.7
29
84.6
24
51.1
29
75.1
28
CoT-VLA
87.5
23
91.6
23
87.6
22
69.0
22
81.1
23
WorldVLA (8) (256 × 256)
85.6
25
89.0
26
82.6
26
59.0
25
79.1
24
WorldVLA (8) (512 × 512)
87.6
22
96.2
18
83.4
25
60.0
24
81.8
22
Table 1: Comparison on the LIBERO benchmark. * represents that the LLM backbone is frozen during training. Our proposed approach is trained on the LIBERO dataset. All metrics are average success rates (%). The best results are highlighted in bold . △ denotes IG-VLA with Latent Spatiotemporal Reasoning replaced by Scene Gist Memory.
Methods
Camera
Robot
Language
Light
Background
Noise
Layout
Avg.
π0⋄
79.6
21.1
72.5
84.7
86.2
68.3
69.4
67.4
π0.5⋄
70.3
41.7
81.1
97.3
94.6
71.8
84.9
75.7
ACoT-VLA ⋄
91.2
62.5
80.3
95.1
91.5
88.3
84.9
84.1
ACoT-VLA
96.6
70.4
79.7
95.1
97.1
95.9
85.0
88.0
IG-VLA (Ours)
95.4
68.9
85.6
97.5
97.2
94.9
85.5
88.5
IG-VLA (Ours) △
95.3
69.0
85.1
97.4
97.5
94.7
84.6
88.3
Table 2: Supervised fine-tuning results on LIBERO-Plus. ⋄ denotes a pipeline with a frozen LLM backbone, and △ denotes IG-VLA with Latent Spatiotemporal Reasoning replaced by Scene Gist Memory. Best results are in bold .
Methods
Guidance
In-dist.
Category
Commonsense
Instruction
Texture
Avg.
IS ↑
PS ↑
IS ↑
PS ↑
IS ↑
PS ↑
IS ↑
PS ↑
IS ↑
PS ↑
IS ↑
PS ↑
π0⋄
Linguistics
67.8
62.7
44.0
33.6
54.9
43.0
58.0
38.7
50.6
42.5
55.0
44.1
π0.5⋄
Linguistics
75.0
60.8
49.6
35.3
57.5
41.6
57.1
30.3
62.0
47.4
60.2
43.1
ACoT-VLA
Action
79.8
66.1
54.1
38.9
52.3
37.8
56.8
39.6
74.6
54.6
63.5
47.4
IG-VLA (Ours)
Action
80.80
61.64
54.20
32.21
58.20
38.97
65.20
48.49
74.99
49.23
66.68
46.11
IG-VLA (Ours) △
Action
80.48
61.23
54.35
31.75
57.82
39.18
64.76
48.22
74.71
48.88
66.26
45.68
Table 3: Comparison on VLABench. IS and PS denote Intention Score and Progress Score. ⋄ denotes pipeline with a frozen LLM backbone, and △ denotes IG-VLA with Latent Spatiotemporal Reasoning replaced by Scene Gist Memory. Best results are in bold .
Table 7
Figure 4: GPU time breakdown(batch size 1, BF16), averaged over three profiled forward passes.
Figure 9
Figure 7: Real-world evaluation on the dual-arm UR3. (a) Pick-and-place-a-screwdriver task execution. (b) Success rate comparison with baselines over 16 trials per method.
Department of Automation, BNRist, Tsinghua University, Beijing, China · Department of Computer Science, The University of Hong Kong, Hong Kong, China · Dexmal, Beijing, China +1