Cross-view spatial reasoning requires a model to align different viewpoints into a coherent spatial representation, yet this ability remains challenging for vision-language models despite being natural to humans. Existing methods typically improve spatial reasoning by updating model weights, which keeps the acquired knowledge implicit and tied to a specific backbone. We propose \textit{SpatialSkill}, a weight-update-free framework that enables a frozen vision-language model to accumulate explicit natural-language reasoning skills from offline trajectories. Unlike symbolic tasks, perceptual skills cannot be reliably verified simply by executing them: a plausible spatial rule may lack visual support or require transformations that the frozen model cannot perform. SpatialSkill therefore admits candidate skills only after visual-grounding and executability checks, constrains manual evolution to prevent harmful regressions, and routes skills by spatial-reasoning category to reduce negative transfer. On CityCube, across four frozen executors, SpatialSkill yields consistent gains, and a 9B executor equipped with SpatialSkill surpasses the strongest closed-source reference in our evaluation. The skills are stored in a versioned natural-language manual, making the reasoning strategies explicit and auditable without modifying model parameters. Code at https://github.com/vindahi/SpatialSkill.
Figures & tables
Figure 1: A perspective-taking (PT) instance and a spatial-relation (SR) instance, each pairing aerial and ground views. Frozen executor fails on both without manuals, while a visually grounded rule from the evolved SpatialSkill manual enables it to correct its prediction without any weight update.
Figure 2: Overview of Self-Evolving Skills for Cross-View Spatial Reasoning (SpatialSkill)
Method
Accuracy (%) ↑
Option-wise Acc. Range (pp) ↓
Overall
Δ
CR
MR
PT
SR
WK
Human
88.3
–
93.1
92.4
87.4
90.2
78.6
–
Paired executors
Qwen3.5-4B + SpatialSkill
52.08
+4.92
43.24
51.22
41.77
56.10
64.06
23.52
Qwen3.5-4B
47.16
–
51.35
53.66
39.87
45.53
52.34
27.28
Qwen3.5-9B + SpatialSkill
57.58
+3.03
59.46
67.07
51.90
52.85
62.50
23.10
Table 1: Main results on the CityCube dataset. Accuracy is reported in percent. Option-wise Acc. Range (pp) is the maximum difference in conditional accuracy across predicted answer options.
Condition
Qwen
Gemma
Direct baseline
47.16
49.24
CoT baseline
49.62
48.67
w/o attribution-aware reflection
50.00
47.35
w/o visual-evidence criterion
50.76
50.74
w/o executability criterion
51.14
51.32
w/o joint admission gate
50.19
50.95
Table 2: Overall ablations validation accuracy (%) on Qwen3.5-4B and Gemma-3-4B.
Table 3: Routing summary and paired error turnover on the held-out validation split.
Method
Overall Acc. (%) ↑
Option-wise Acc. Range (pp) ↓
Qwen3.5-4B Base
29.30
10.45
Qwen3.5-4B Manual + route
31.30
8.01
GPT-5.6-luna Base
34.30
10.62
GPT-5.6-luna Manual + route
38.40
3.95
Table 4: Own-base and frozen manual-routed MMSI-Bench accuracy and option-wise accuracy range. Each executor’s frozen manual and source-derived routing policy is applied to the 1,000-example transfer set .
Reflection Model
Executor
Accuracy (%) ↑
Δ (pp) ↑
None
Qwen3.5-4B
47.16
–
Qwen3-4B
Qwen3.5-4B
50.57
+3.41
Qwen3-8B
Qwen3.5-4B
49.05
+1.89
Qwen3.5-9B
Qwen3.5-4B
52.08
+4.92
Table 5: Reflection-model sensitivity with Qwen3.5-4B as the frozen executor. Only the reflection model is changed; all other evolution and inference settings are fixed.
National Key Laboratory for Novel Software Technology, Nanjing University, China · School of Artificial Intelligence, Nanjing University, China · College of Computer Science, Sichuan University
Department of Computer Science, Wayne State University · Department of Radiation Oncology, Henry Ford Health · Department of Electrical and Computer Engineering, The Ohio State University +1