cs.CVSep 28, 2026

SpatialSkill: Self-Evolving Skills for Cross-View Spatial Reasoning

Authors: Ruifan Zuo, Guocheng Hu, Wanshui Gan, Junyi Wang, Xiang Lei, Tian Gan

Organizations: Shandong University · Shanghai AI Laboratory · Zhiyang Innovation Co., Ltd.

Abstract

Cross-view spatial reasoning requires a model to align different viewpoints into a coherent spatial representation, yet this ability remains challenging for vision-language models despite being natural to humans. Existing methods typically improve spatial reasoning by updating model weights, which keeps the acquired knowledge implicit and tied to a specific backbone. We propose \textit{SpatialSkill}, a weight-update-free framework that enables a frozen vision-language model to accumulate explicit natural-language reasoning skills from offline trajectories. Unlike symbolic tasks, perceptual skills cannot be reliably verified simply by executing them: a plausible spatial rule may lack visual support or require transformations that the frozen model cannot perform. SpatialSkill therefore admits candidate skills only after visual-grounding and executability checks, constrains manual evolution to prevent harmful regressions, and routes skills by spatial-reasoning category to reduce negative transfer. On CityCube, across four frozen executors, SpatialSkill yields consistent gains, and a 9B executor equipped with SpatialSkill surpasses the strongest closed-source reference in our evaluation. The skills are stored in a versioned natural-language manual, making the reasoning strategies explicit and auditable without modifying model parameters. Code at https://github.com/vindahi/SpatialSkill.

Figures & tables

Explore similar work

CardsList
  1. Self-Evolving Neuro-Symbolic Skills for Tool-Augmented Spatial Reasoning

    Aug 8, 2026Shi-Yu Tian, Zhuo-Xia Wang, Xuan-Yi Zhu +6Spatial ReasoningNeuro-Symbolic Framework

  2. SpatialCORE: Confidence-Aware Grounded Spatial Reasoning in Large Vision--Language Models

    Sep 30, 2026Rafi Ibn Sultan, Xiangyu Zhou, Md. Sajid Alam Chowdhury +4Spatial ReasoningRecent Vision-Language Models

  3. SpatialThinker: Reinforcing Scene Graph-Grounded Spatial Reasoning via Dense Rewards

    Nov 10, 2025Hunar Batra, Haoqin Tu, Hardy Chen +3Spatial ReasoningSpatial Grounding