Cross-view spatial reasoning requires a model to align different viewpoints into a coherent spatial representation, yet this ability remains challenging for vision-language models despite being natural to humans. Existing methods typically improve spatial reasoning by updating model weights, which keeps the acquired knowledge implicit and tied to a specific backbone. We propose \textit{SpatialSkill}, a weight-update-free framework that enables a frozen vision-language model to accumulate explicit natural-language reasoning skills from offline trajectories. Unlike symbolic tasks, perceptual skills cannot be reliably verified simply by executing them: a plausible spatial rule may lack visual support or require transformations that the frozen model cannot perform. SpatialSkill therefore admits candidate skills only after visual-grounding and executability checks, constrains manual evolution to prevent harmful regressions, and routes skills by spatial-reasoning category to reduce negative transfer. On CityCube, across four frozen executors, SpatialSkill yields consistent gains, and a 9B executor equipped with SpatialSkill surpasses the strongest closed-source reference in our evaluation. The skills are stored in a versioned natural-language manual, making the reasoning strategies explicit and auditable without modifying model parameters. Code at https://github.com/vindahi/SpatialSkill.
Figures & tables
Figure 1: A perspective-taking (PT) instance and a spatial-relation (SR) instance, each pairing aerial and ground views. Frozen executor fails on both without manuals, while a visually grounded rule from the evolved SpatialSkill manual enables it to correct its prediction without any weight update.
Figure 2: Overview of Self-Evolving Skills for Cross-View Spatial Reasoning (SpatialSkill)
Method
Accuracy (%) ↑
Option-wise Acc. Range (pp) ↓
Overall
Δ
CR
MR
PT
SR
WK
Human
88.3
–
93.1
92.4
87.4
90.2
78.6
–
Paired executors
Qwen3.5-4B + SpatialSkill
52.08
+4.92
43.24
51.22
41.77
56.10
64.06
23.52
Qwen3.5-4B
47.16
–
51.35
53.66
39.87
45.53
52.34
27.28
Qwen3.5-9B + SpatialSkill
57.58
+3.03
59.46
67.07
51.90
52.85
62.50
23.10
Table 1: Main results on the CityCube dataset. Accuracy is reported in percent. Option-wise Acc. Range (pp) is the maximum difference in conditional accuracy across predicted answer options.
Condition
Qwen
Gemma
Direct baseline
47.16
49.24
CoT baseline
49.62
48.67
w/o attribution-aware reflection
50.00
47.35
w/o visual-evidence criterion
50.76
50.74
w/o executability criterion
51.14
51.32
w/o joint admission gate
50.19
50.95
Table 2: Overall ablations validation accuracy (%) on Qwen3.5-4B and Gemma-3-4B.
Table 3: Routing summary and paired error turnover on the held-out validation split.
Method
Overall Acc. (%) ↑
Option-wise Acc. Range (pp) ↓
Qwen3.5-4B Base
29.30
10.45
Qwen3.5-4B Manual + route
31.30
8.01
GPT-5.6-luna Base
34.30
10.62
GPT-5.6-luna Manual + route
38.40
3.95
Table 4: Own-base and frozen manual-routed MMSI-Bench accuracy and option-wise accuracy range. Each executor’s frozen manual and source-derived routing policy is applied to the 1,000-example transfer set .
Reflection Model
Executor
Accuracy (%) ↑
Δ (pp) ↑
None
Qwen3.5-4B
47.16
–
Qwen3-4B
Qwen3.5-4B
50.57
+3.41
Qwen3-8B
Qwen3.5-4B
49.05
+1.89
Qwen3.5-9B
Qwen3.5-4B
52.08
+4.92
Table 5: Reflection-model sensitivity with Qwen3.5-4B as the frozen executor. Only the reflection model is changed; all other evolution and inference settings are fixed.
Large vision-language models have achieved strong performance in multimodal reasoning, but they remain unreliable on fine-grained spatial tasks that demand both precise spatial perception and fine-grained geometric computation beyond end-to-end generation. Tool augmentation offers a natural solution, while existing methods either plan tool calls from scratch without explicit dependency constraints or rely on fixed pipelines that are redundant and generalize poorly across spatial tasks. An effective spatial reasoning agent should instead accumulate reusable experience and adaptively compose it for new problems. To this end, we propose NeSy-Spatial, a neuro-symbolic framework for self-evolving spatial skills. NeSy-Spatial abstracts tool interactions and geometric operations into typed executable atomic instructions and composes them into two complementary skill types: Tool-Use Skills for organizing tool execution and Geometry Skills for structured geometric reasoning. During inference, NeSy-Spatial retrieves and executes relevant skills in a closed-loop process. During evolution, it analyzes buffered successful and failed trajectories to refine skill structures and prune unreliable or inactive entries. Experiments on three spatial reasoning benchmarks show that NeSy-Spatial consistently improves reasoning accuracy with more precise tool utilization.
Shi-Yu Tian, Zhuo-Xia Wang, Xuan-Yi Zhu +6
National Key Laboratory for Novel Software Technology, Nanjing University, China · School of Artificial Intelligence, Nanjing University, China · College of Computer Science, Sichuan University
Large Vision-Language Models (LVLMs) have made remarkable progress across visual perception tasks, yet spatial reasoning remains a persistent weakness, especially for questions that require reasoning over visual space. Recent spatial-reasoning methods incorporate generated grounding, where models predict bounding boxes, masks, or other localization outputs for task-relevant objects as part of their reasoning trace. However, these approaches typically optimize final-answer correctness alone, allowing correct answers to be rewarded even when the model does not reason from confidently localized task-relevant objects. We introduce SpatialCORE (Spatially COnfident REasoning), a post-training framework that turns the model's own confidence in generated grounding into a learning signal for spatial reasoning. Its central idea is to reinforce grounding that is both accurate and confident, encouraging the model to reason from confidently localized task-relevant objects. SpatialCORE realizes this through a self-regulating spatial reward that weights each predicted bounding box's matching quality by its coordinate-token confidence. An answer gate further ties grounding optimization to final-answer correctness. SpatialCORE achieves state-of-the-art results among open-source and specialized spatial reasoning models across diverse benchmarks, and transfers effectively in zero-shot settings to unseen data distributions. The source code is available at https://github.com/rafiibnsultan/SpatialCORE.
Rafi Ibn Sultan, Xiangyu Zhou, Md. Sajid Alam Chowdhury +4
Department of Computer Science, Wayne State University · Department of Radiation Oncology, Henry Ford Health · Department of Electrical and Computer Engineering, The Ohio State University +1
Multimodal large language models (MLLMs) have achieved remarkable progress in vision-language tasks, but continue to struggle with spatial reasoning. Existing spatial MLLMs rely on large-scale datasets, explicit 3D inputs, architecture-specific modifications, or sparse Reinforcement Learning (RL) methods that provide insufficient guidance for spatially-grounded reasoning. We introduce SpatialThinker. To our knowledge, it is the first MLLM unifying Scene Graph Generation (SGG) and visual reasoning in a single pass via online RL. The model simulates human-like spatial perception by constructing a mental scene graph of task-relevant objects and relations, and reasoning toward an answer via dense spatial rewards. Our contributions are threefold: (1) SGG-grounded reasoning: integrating SGG directly within the reasoning chain rather than as a disjoint preprocessing step; (2) STVQA-7K: a high-quality spatial VQA training dataset via a scalable synthesis pipeline; and (3) a dense spatial reward design that enforces structured grounding during RL and generalizes to improve broad visual perception. SpatialThinker-7B achieves 3.6× larger gains over SFT and 1.7× better in- and out-of-distribution generalization than sparse RL. Trained on only 7K samples, SpatialThinker-7B matches GPT-5 and outperforms GPT-4o, while SpatialThinker-30B surpasses both GPT-5 and Claude 4 Sonnet on average across 14 spatial and real-world benchmarks, demonstrating that structured spatial grounding with reward-aligned reasoning enables robust spatial understanding with limited data.
Hunar Batra, Haoqin Tu, Hardy Chen +3
University of Oxford · University of California, Santa Cruz