Humans often solve spatial problems by mentally simulating visual transformations. In contrast, conventional vision-language models (VLMs) reason primarily through language. We investigate whether VLMs can solve spatial problems by reasoning with both text and generated visual states. To this end, we introduce WM-VLM, which equips a pretrained VLM with a lightweight world model branch for generating intermediate visual states. Our two-stage training first teaches the model to generate the next visual state and then to use that state for reasoning. We programmatically construct spatial reasoning tasks with verifiable intermediate visual states. These tasks allow us to evaluate how well the model generates visual states and how much it relies on them to answer the question. On 2D and 3D mental rotation tasks, WM-VLM consistently outperforms the supervised fine-tuned backbone, with gains of up to 39.25 percentage points. Ablations suggest that these gains depend on the generated visual states, as removing or corrupting them sharply reduces performance. Together, these results suggest that internal world models offer a promising path toward VLMs that reason in both language and visual space.
Figures & tables
Figure 1: The internal world model module and training objective positively contributes to solving spatial-related problems, which finally helps outperform the finetuned base VLM. WM-VLM exhibits a sudden performance transition at around step 25k, then plateaued, yielding a 39.25 points gain on ID set and 28.8 points gain on OOD set. We train WM-VLM with the middle-layer configuration and evaluate each checkpoint on Tetris-2D-ID (left) and Tetris-2D-OOD (right) splits. Training step indicates the training steps taken in Stage 1. All data points are obtained following an additional Stage 2 training. Each panel compares the accuracy of WM-VLM and Qwen2.5-VL-7B SFT (left axis). We also show the cosine similarity between the generated visual states and ground truth embedding during training (right axis, in purple). The dashed horizontal line indicates random-guess accuracy (25%).
Figure 2: WM-VLM performs interleaved visual-textual reasoning by alternating between textual mental actions and generated visual states, before producing the final answer.
Figure 3: Architecture of WM-VLM , instantiated with our light Mixture-of-Transformers (light MoT) design. The understanding branch, comprising the vision encoder and language decoder, is initialized from a pretrained VLM. The shallow generation branch contains k layers, each paired with a corresponding layer in a consecutive k -layer block of the understanding branch. Text and clean visual tokens are routed to the understanding branch, whereas noisy visual tokens are routed to the generation branch. All tokens interact through global self-attention.
Method
Visual Reasoning Type
Objective
Tetris-2D
Tetris-3D
ID
OOD
SC-ID
C-ID
OOD
Qwen2.5-VL-7B-Inst.
(Text only)
AR-CE
22.00
23.80
21.25
23.00
21.40
+ SFT
(Text only)
AR-CE
48.25
46.20
71.75
44.50
55.50
LatentUM
Discrete Visual Tokens
AR-CE
29.25
23.00
58.50
37.75
39.00
Mirage
Continuous Visual Tokens
AR-Cosine
47.75
43.00
66.25
46.50
49.80
ThinkMorph
Generated Image
FM-MSE
22.50
21.60
0.00*
0.00*
0.00*
Table 1: Performance comparison of methods with different reasoning types and training objectives. All models in this table are fine-tuned on Tetris-2D or Tetris-3D, except Qwen2.5-VL-7B-Inst. Results are accuracy (%); Δ denotes the absolute improvement of WM-VLM over Qwen2.5-VL-7B-Instruct SFT in percentage points. AR, FM, CE, and MSE denote autoregressive, flow matching, cross-entropy, and mean squared error, respectively. *Fine-tuned BAGEL (ThinkMorph) fails to emit images and generate answers on Tetris-3D, resulting in 0% accuracy.
Metric
Tetris-2D-ID
Tetris-2D-OOD
Tetris-3D-SC-ID
Tetris-3D-C-ID
Tetris-3D-OOD
Cosine
0.9347
0.7523
0.8511
0.7921
0.7746
Retrieval top-1
96.50
53.00
65.75
20.75
13.40
Standard accuracy
87.50
75.00
91.00
66.25
71.80
Accuracy-✓
89.38
79.25
99.24
92.77
86.57
Δ
+1.88
+4.25
+8.24
+26.52
+14.77
Accuracy-✗
35.71
70.21
75.18
59.31
69.52
Table 2: Relationship between visual-token retrieval quality and answer accuracy. Cosine denotes the average cosine similarity between the generated and ground truth visual embeddings. Retrieval top-1 denotes the percentage of times the correct ground truth embedding is retrieved. Accuracy-✓ is the accuracy when the retrieval is correct, while Accuracy-✗ is the accuracy when the retrieval is wrong. Δ denotes the absolute change relative to standard accuracy. The ϕ coefficient is computed between two binary indicators: whether the correct ground-truth embedding is retrieved and whether the final answer is correct.
Metric
all-white pixel
shuffle 19×19 patches
Default
Tetris-2D-ID
49.25
48.25
47.25
Δ
+2.00
+1.00
–
Tetris-2D-OOD
46.60
46.20
46.20
Δ
+0.40
0.00
–
Table 3: Ablations on alternating generated image pixels on Tetris-2D. We report accuracy (%). Δ denotes the change relative to standard inference with generated pixels.
Metric
all-white pixel
shuffle 19×19 patches
Default
Tetris-2D-ID
49.25
48.25
47.25
Δ
+2.00
+1.00
–
Tetris-2D-OOD
46.60
46.20
46.20
Δ
+0.40
0.00
–
Table 3: Ablations on alternating generated image pixels on Tetris-2D. We report accuracy (%). Δ denotes the change relative to standard inference with generated pixels.
Metric
all-zero embedding
randomized tokens
shuffle visual tokens
Default
Tetris-2D-ID
14.00
23.50
81.50
87.50
Δ
-73.50
-64.00
-6.00
–
Tetris-2D-OOD
15.40
20.80
72.60
75.00
Δ
-59.60
-54.20
-2.40
–
Table 4: Ablations on alternating generated visual tokens onTetris-2D. Results are accuracy (%); Δ denotes the change relative to evaluation with the default setup.
Metric
all-zero embedding
randomized tokens
shuffle visual tokens
default
Tetris-3D-SC-ID
15.50
23.25
81.00
91.00
Δ
-75.50
-67.75
-10.00
–
Tetris-3D-C-ID
16.00
25.00
60.25
66.25
Δ
-50.25
-41.25
-6.00
–
Tetris-3D-OOD
14.40
20.40
64.20
71.80
Δ
-57.40
-51.40
-7.60
–
Table 5: Ablations on alternating generated visual tokens on Tetris-3D. Results are accuracy (%); Δ denotes the change relative to evaluation with the default setup.
Figure 4: Effect of visual-token count on model performance. We report final-step MSE, training-time token similarity, and performance on Tetris-2D-ID and Tetris-2D-OOD. MSE compares the generated and ground-truth visual embeddings. All-token cosine is computed after concatenating all tokens in each block, whereas per-token cosine averages the similarities between corresponding generated and ground-truth tokens.
Figure 5: Effect of the ratio between cross-entropy and flow-matching losses on model performance.
Figure 6: Training curves for joint training with loss ratio α=0.1 and β=0.9 .
Figure 7: Layer-wise analysis of visual latent quality and downstream performance in the VLM.
Figure 8: Effect of the layer-window selection on model performance. For example, ”8-11” means the generation branch is connected to the understanding branch from layer 8 to layer 11.
Appendix figures & tables6 assets
Supplementary material from the paper’s appendix.
Appendix
Figure 9: Representative samples from Tetris-2D and Tetris-3D. Text denotes mental rotation actions and images denote their resulting visual states. The 2D example stores the final state after two 90∘ rotations, while the 3D example explicitly interleaves each 90∘ rotation with its corresponding visual state.
Model
Think Mode
Eval Setting
Tetris-2D
Tetris-3D
VSP-Nav
ID
OOD
ID
OOD
Qwen2.5-VL-7B-Instruct
✗
Question only
22.00
23.80
23.00
21.40
6.50
Q + RT
23.50
23.20
20.25
21.00
21.50
Q + RT + I
26.25
26.20
60.50
59.60
24.83
Qwen3-VL-8B-Instruct
✗
Question only
23.00
21.80
9.00
9.20
1.33
Q + RT
20.50
21.00
7.00
7.40
0.67
Appendix
Table 6: Evaluation results of VLMs on the test sets of Tetris-2D, Tetris-3D, and VSP-Nav under different reasoning-input settings. Think Mode is the model’s built-in capability from their pre-training. Qwen2.5-VL-7B-Instruct and Qwen3-VL-8B-Instruct do not support think mode by default. Qwen3.5-9B, Qwen3.6-27B, and Qwen3.7-40B support think mode. We test these models in both modes. ✗ indicates the think mode is disabled; ✓ indicates the think mode is enabled. Q denotes the question input, RT denotes reasoning text, and I denotes intermediate reasoning images. When testing with the Q + RT setting, we replace the intermediate reasoning images with token ”[IMAGINATION]”.
Training Data
Method
WM Objective
Accuracy
3×3
4×4
5×5
6×6
7×7
8×8
Proprietary interleaved data + ThinkMorph-SN
ThinkMorph (BAGEL)
✓
82.50
93
95
94
78
74
61
/
Qwen2.5-VL-7B-Inst.
✗
6.50
23
4
5
3
2
2
ThinkMorph-SN
Qwen2.5-VL-7B-Inst. SFT
✗
43.17
79
65
50
34
17
14
LatentUM
✓
52.00
90
75
59
42
24
22
Mirage
✓
42.50
79
58
56
29
21
12
WM-VLM (Ours)
✓
50.00
98
83
63
31
18
7
Appendix
Table 7: Maze-navigation results across different training datasets and grid sizes. Accuracy and per-grid-size results are reported in percent.
Figure 10: Effect of the number of middle layers on model performance.
Hyperparameter
Stage 1
Stage 2
Initialization
Qwen2.5-VL-7B-Instruct
Stage 1 checkpoint
Learning rate
1e-4
1e-5
Optimizer
AdamW
AdamW
LR scheduler
Constant
Cosine
Warmup ratio
0
0.03
Weight decay
0
0.01
Appendix
Table 8: Hyperparameters for the two-stage training procedure on Tetris-2D.
Figure 11: Effect of generation-branch initialization on downstream performance across different insertion depths.
Institute of Automation, Chinese Academy of Sciences · School of Artificial Intelligence, University of Chinese Academy of Sciences · Harbin Institute of Technology +2