Humans often solve spatial problems by mentally simulating visual transformations. In contrast, conventional vision-language models (VLMs) reason primarily through language. We investigate whether VLMs can solve spatial problems by reasoning with both text and generated visual states. To this end, we introduce WM-VLM, which equips a pretrained VLM with a lightweight world model branch for generating intermediate visual states. Our two-stage training first teaches the model to generate the next visual state and then to use that state for reasoning. We programmatically construct spatial reasoning tasks with verifiable intermediate visual states. These tasks allow us to evaluate how well the model generates visual states and how much it relies on them to answer the question. On 2D and 3D mental rotation tasks, WM-VLM consistently outperforms the supervised fine-tuned backbone, with gains of up to 39.25 percentage points. Ablations suggest that these gains depend on the generated visual states, as removing or corrupting them sharply reduces performance. Together, these results suggest that internal world models offer a promising path toward VLMs that reason in both language and visual space.
Figures & tables
Figure 1: The internal world model module and training objective positively contributes to solving spatial-related problems, which finally helps outperform the finetuned base VLM. WM-VLM exhibits a sudden performance transition at around step 25k, then plateaued, yielding a 39.25 points gain on ID set and 28.8 points gain on OOD set. We train WM-VLM with the middle-layer configuration and evaluate each checkpoint on Tetris-2D-ID (left) and Tetris-2D-OOD (right) splits. Training step indicates the training steps taken in Stage 1. All data points are obtained following an additional Stage 2 training. Each panel compares the accuracy of WM-VLM and Qwen2.5-VL-7B SFT (left axis). We also show the cosine similarity between the generated visual states and ground truth embedding during training (right axis, in purple). The dashed horizontal line indicates random-guess accuracy (25%).
Figure 2: WM-VLM performs interleaved visual-textual reasoning by alternating between textual mental actions and generated visual states, before producing the final answer.
Figure 3: Architecture of WM-VLM , instantiated with our light Mixture-of-Transformers (light MoT) design. The understanding branch, comprising the vision encoder and language decoder, is initialized from a pretrained VLM. The shallow generation branch contains k layers, each paired with a corresponding layer in a consecutive k -layer block of the understanding branch. Text and clean visual tokens are routed to the understanding branch, whereas noisy visual tokens are routed to the generation branch. All tokens interact through global self-attention.
Method
Visual Reasoning Type
Objective
Tetris-2D
Tetris-3D
ID
OOD
SC-ID
C-ID
OOD
Qwen2.5-VL-7B-Inst.
(Text only)
AR-CE
22.00
23.80
21.25
23.00
21.40
+ SFT
(Text only)
AR-CE
48.25
46.20
71.75
44.50
55.50
LatentUM
Discrete Visual Tokens
AR-CE
29.25
23.00
58.50
37.75
39.00
Mirage
Continuous Visual Tokens
AR-Cosine
47.75
43.00
66.25
46.50
49.80
ThinkMorph
Generated Image
FM-MSE
22.50
21.60
0.00*
0.00*
0.00*
Table 1: Performance comparison of methods with different reasoning types and training objectives. All models in this table are fine-tuned on Tetris-2D or Tetris-3D, except Qwen2.5-VL-7B-Inst. Results are accuracy (%); Δ denotes the absolute improvement of WM-VLM over Qwen2.5-VL-7B-Instruct SFT in percentage points. AR, FM, CE, and MSE denote autoregressive, flow matching, cross-entropy, and mean squared error, respectively. *Fine-tuned BAGEL (ThinkMorph) fails to emit images and generate answers on Tetris-3D, resulting in 0% accuracy.
Metric
Tetris-2D-ID
Tetris-2D-OOD
Tetris-3D-SC-ID
Tetris-3D-C-ID
Tetris-3D-OOD
Cosine
0.9347
0.7523
0.8511
0.7921
0.7746
Retrieval top-1
96.50
53.00
65.75
20.75
13.40
Standard accuracy
87.50
75.00
91.00
66.25
71.80
Accuracy-✓
89.38
79.25
99.24
92.77
86.57
Δ
+1.88
+4.25
+8.24
+26.52
+14.77
Accuracy-✗
35.71
70.21
75.18
59.31
69.52
Table 2: Relationship between visual-token retrieval quality and answer accuracy. Cosine denotes the average cosine similarity between the generated and ground truth visual embeddings. Retrieval top-1 denotes the percentage of times the correct ground truth embedding is retrieved. Accuracy-✓ is the accuracy when the retrieval is correct, while Accuracy-✗ is the accuracy when the retrieval is wrong. Δ denotes the absolute change relative to standard accuracy. The ϕ coefficient is computed between two binary indicators: whether the correct ground-truth embedding is retrieved and whether the final answer is correct.
Metric
all-white pixel
shuffle 19×19 patches
Default
Tetris-2D-ID
49.25
48.25
47.25
Δ
+2.00
+1.00
–
Tetris-2D-OOD
46.60
46.20
46.20
Δ
+0.40
0.00
–
Table 3: Ablations on alternating generated image pixels on Tetris-2D. We report accuracy (%). Δ denotes the change relative to standard inference with generated pixels.
Metric
all-white pixel
shuffle 19×19 patches
Default
Tetris-2D-ID
49.25
48.25
47.25
Δ
+2.00
+1.00
–
Tetris-2D-OOD
46.60
46.20
46.20
Δ
+0.40
0.00
–
Table 3: Ablations on alternating generated image pixels on Tetris-2D. We report accuracy (%). Δ denotes the change relative to standard inference with generated pixels.
Metric
all-zero embedding
randomized tokens
shuffle visual tokens
Default
Tetris-2D-ID
14.00
23.50
81.50
87.50
Δ
-73.50
-64.00
-6.00
–
Tetris-2D-OOD
15.40
20.80
72.60
75.00
Δ
-59.60
-54.20
-2.40
–
Table 4: Ablations on alternating generated visual tokens onTetris-2D. Results are accuracy (%); Δ denotes the change relative to evaluation with the default setup.
Metric
all-zero embedding
randomized tokens
shuffle visual tokens
default
Tetris-3D-SC-ID
15.50
23.25
81.00
91.00
Δ
-75.50
-67.75
-10.00
–
Tetris-3D-C-ID
16.00
25.00
60.25
66.25
Δ
-50.25
-41.25
-6.00
–
Tetris-3D-OOD
14.40
20.40
64.20
71.80
Δ
-57.40
-51.40
-7.60
–
Table 5: Ablations on alternating generated visual tokens on Tetris-3D. Results are accuracy (%); Δ denotes the change relative to evaluation with the default setup.
Figure 4: Effect of visual-token count on model performance. We report final-step MSE, training-time token similarity, and performance on Tetris-2D-ID and Tetris-2D-OOD. MSE compares the generated and ground-truth visual embeddings. All-token cosine is computed after concatenating all tokens in each block, whereas per-token cosine averages the similarities between corresponding generated and ground-truth tokens.
Figure 5: Effect of the ratio between cross-entropy and flow-matching losses on model performance.
Figure 6: Training curves for joint training with loss ratio α=0.1 and β=0.9 .
Figure 7: Layer-wise analysis of visual latent quality and downstream performance in the VLM.
Figure 8: Effect of the layer-window selection on model performance. For example, ”8-11” means the generation branch is connected to the understanding branch from layer 8 to layer 11.
Appendix figures & tables6 assets
Supplementary material from the paper’s appendix.
Appendix
Figure 9: Representative samples from Tetris-2D and Tetris-3D. Text denotes mental rotation actions and images denote their resulting visual states. The 2D example stores the final state after two 90∘ rotations, while the 3D example explicitly interleaves each 90∘ rotation with its corresponding visual state.
Model
Think Mode
Eval Setting
Tetris-2D
Tetris-3D
VSP-Nav
ID
OOD
ID
OOD
Qwen2.5-VL-7B-Instruct
✗
Question only
22.00
23.80
23.00
21.40
6.50
Q + RT
23.50
23.20
20.25
21.00
21.50
Q + RT + I
26.25
26.20
60.50
59.60
24.83
Qwen3-VL-8B-Instruct
✗
Question only
23.00
21.80
9.00
9.20
1.33
Q + RT
20.50
21.00
7.00
7.40
0.67
Appendix
Table 6: Evaluation results of VLMs on the test sets of Tetris-2D, Tetris-3D, and VSP-Nav under different reasoning-input settings. Think Mode is the model’s built-in capability from their pre-training. Qwen2.5-VL-7B-Instruct and Qwen3-VL-8B-Instruct do not support think mode by default. Qwen3.5-9B, Qwen3.6-27B, and Qwen3.7-40B support think mode. We test these models in both modes. ✗ indicates the think mode is disabled; ✓ indicates the think mode is enabled. Q denotes the question input, RT denotes reasoning text, and I denotes intermediate reasoning images. When testing with the Q + RT setting, we replace the intermediate reasoning images with token ”[IMAGINATION]”.
Training Data
Method
WM Objective
Accuracy
3×3
4×4
5×5
6×6
7×7
8×8
Proprietary interleaved data + ThinkMorph-SN
ThinkMorph (BAGEL)
✓
82.50
93
95
94
78
74
61
/
Qwen2.5-VL-7B-Inst.
✗
6.50
23
4
5
3
2
2
ThinkMorph-SN
Qwen2.5-VL-7B-Inst. SFT
✗
43.17
79
65
50
34
17
14
LatentUM
✓
52.00
90
75
59
42
24
22
Mirage
✓
42.50
79
58
56
29
21
12
WM-VLM (Ours)
✓
50.00
98
83
63
31
18
7
Appendix
Table 7: Maze-navigation results across different training datasets and grid sizes. Accuracy and per-grid-size results are reported in percent.
Figure 10: Effect of the number of middle layers on model performance.
Hyperparameter
Stage 1
Stage 2
Initialization
Qwen2.5-VL-7B-Instruct
Stage 1 checkpoint
Learning rate
1e-4
1e-5
Optimizer
AdamW
AdamW
LR scheduler
Constant
Cosine
Warmup ratio
0
0.03
Weight decay
0
0.01
Appendix
Table 8: Hyperparameters for the two-stage training procedure on Tetris-2D.
Figure 11: Effect of generation-branch initialization on downstream performance across different insertion depths.
Vision-language models (VLMs) have shown strong performance on static visual understanding, yet they still struggle with dynamic spatial reasoning that requires imagining how scenes evolve under egocentric motion. Recent efforts address this limitation either by scaling spatial supervision with synthetic data or by coupling VLMs with world models at inference time. However, the former often lacks explicit modeling of motion-conditioned state transitions, while the latter incurs substantial computational overhead. In this work, we propose World2VLM, a training framework that distills spatial imagination from a generative world model into a vision-language model. Given an initial observation and a parameterized camera trajectory, we use a view-consistent world model to synthesize geometrically aligned future views and derive structured supervision for both forward (action-to-outcome) and inverse (outcome-to-action) spatial reasoning. We post-train the VLM with a two-stage recipe on a compact dataset generated by this pipeline and evaluate it on multiple spatial reasoning benchmarks. World2VLM delivers consistent improvements over the base model across diverse benchmarks, including SAT-Real, SAT-Synthesized, VSI-Bench, and MindCube. It also outperforms the test-time world-model-coupled methods while eliminating the need for expensive inference-time generation. Our results suggest that world models can serve not only as inference-time tools, but also as effective training-time teachers, enabling VLMs to internalize spatial imagination in a scalable and efficient manner.
Wanyue Zhang, Wenxiang Wu, Wang Xu +6
Institute of Automation, Chinese Academy of Sciences · School of Artificial Intelligence, University of Chinese Academy of Sciences · Harbin Institute of Technology +2
Modern Vision-Language Models (VLMs) achieve strong semantic recognition, yet remain brittle on elementary spatial relations such as left of, on, behind, and between. One cause of this failure arises before language reasoning begins: the visual pathway may compress or discard critical 3D structural cues during feature extraction, so the language model receives image representations that are already insufficient for reliable spatial judgment. We introduce GeoWorld-VLM, a VLM-side distillation framework that transfers geometric structure from frozen camera-conditioned video world models into VLMs. GeoWorld-VLM fine-tunes only the image encoder and multimodal projector, aligning post-projector image features with intermediate world-model representations while leaving the main backbone frozen. Given images, a prompt, and a sampled camera trajectory, the world-model teacher converts static visual input into a synthetic multi-view spatial signal. Training combines spatial answer supervision, teacher-student feature alignment, and a preservation anchor to the original VLM. Since the language model remains frozen, GeoWorld-VLM preserves the original model's linguistic capabilities while attributing spatial improvements to the enhanced visual pathway. To evaluate the effectiveness and generality of the proposed method, we apply GeoWorld-VLM to two distinct VLM architectures and observe consistent improvements across both backbones. GeoWorld-VLM improves performance by approximately 4 percent on both the What'sUp and VSR benchmarks, suggesting that world-model-guided visual alignment generalizes across model structures and spatial reasoning datasets.
Renjie Gu, Kaichen Zhou, Yan Luo +1
Harvard AI and Robotics Lab · Kempner Institute for the Study of Natural and Artificial Intelligence · Harvard University
While Vision-Language Models (VLMs) have shown strong visual reasoning capabilities, their spatial reasoning abilities remain largely constrained to the observed images and text-oriented chain-of-thought. They often struggle to infer unobserved layouts, maintain cross-view consistency, and reason from alternative viewpoints when only limited egocentric observations are available. In this work, we study this problem as thinking with imagination, where a VLM actively acquires imagined visual evidence by interacting with a world simulator during reasoning. We propose Astra, an agentic spatial reasoning framework that empowers VLMs with action-conditioned visual imagination. Specifically, Astra couples Astra-VL, an RL-trained VLM policy, with Astra-WM, a Bagel-based world simulator that generates novel-view observations from context images and natural-language camera motions. To provide reliable imagined evidence, Astra-WM is trained with view consistency tuning to improve pose and content consistency across views. In the RL stage, we propose a world-simulator-in-the-loop two-phase RL curriculum to stabilize tool-use exploration and advance the model's ability to invoke the simulator only when imagined observations improve over direct answering. Experiments demonstrate that both the world simulator and the agentic policy are necessary: Astra-WM improves simulator-augmented Gemini-3-Flash on MMSI-Bench from 45.1 to 49.5, while Astra-VL improves the Qwen3-VL backbone from 29.8 to 38.8 on MMSI-Bench and from 36.8 to 42.7 on MindCube. These results show that imagined observations can provide useful spatial evidence, but effective world-model-augmented reasoning requires learning when, where, and how to imagine.
Chenming Zhu, Jingli Lin, Yilin Long +4
1The University of Hong Kong · 2Shanghai AI Laboratory · 3Shanghai Jiao Tong University +2