cs.CVSep 30, 2026

Uruqi: Learning Spatial Cognition from Visual Experience

Authors: Shichao Li, Meiqi Wang, Fei Su, Zhicheng Zhao

Organizations: Beijing University of Posts and Telecommunications · Tsinghua University

Abstract

Spatial intelligence requires maintaining a coherent understanding of the world as the embodied agent moves. Like humans, the agent must use its own motion to interpret changes across observations and update object locations and spatial relations accordingly. Despite spatial post-training having substantially broadened the spatial intelligence of vision-language models (VLMs), they still struggle with two atomic spatial capabilities: tracking self-motion and mapping the surrounding world during motion. To address this gap, we provide dense multi-turn supervision over interleaved atomic capabilities within each training episode, mimicking the visual experience of a continuously moving agent that reasons as it observes. To scale this up, we synthesize 11,738 motif-driven camera trajectories over a broad range of 3D scenes, supporting self-motion tracking, persistent object mapping, and rich spatial operations within each visual experience. By training models to reason over these atomic questions, our URUQISyn_{\mathrm{Syn}}-8B improves accuracy from 15.84% to 47.73% on our Uruqi benchmark comprising 52k questions across 2.7k episodes. URUQI-SI-Mix-8B further reaches 50.41%, comparable to the 50.08% achieved by GPT-6 Astra. Trained solely on our synthesized data, URUQISyn_{\mathrm{Syn}}-8B achieves an average relative accuracy improvement of 17.13% over its InternVL3-8B backbone across three external spatial benchmarks. These results highlight continuous visual experience as a scalable source of supervision for developing spatial cognition in VLMs.

Figures & tables

Appendix figures & tables8 assets

Supplementary material from the paper’s appendix.

Appendix

Explore similar work

CardsList
  1. World2VLM: Distilling World Model Imagination into VLMs for Dynamic Spatial Reasoning

    Apr 29, 2026Wanyue Zhang, Wenxiang Wu, Wang Xu +6Spatial ReasoningVideo World Models

  2. SpatioLM: Towards General Physical Spatial Intelligence in Vision-Language Models

    Aug 3, 2026Jing Wu, Jianhua Wu, Jiayi Guan +5Recent Vision-Language ModelsStable Spatial Understanding

  3. KilometerVision: A New Frontier for Large-Scale Spatial Intelligence in VLMs

    Sep 30, 2026Aravindh Mahendran, Michael King, Matthew Koichi Grimes +12Stable Spatial UnderstandingFrontiers