cs.CVOct 1, 2026

From Reasoning Failures to Composable Video Spatial Intelligence

Authors: Pengzhan Sun, Junbin Xiao, Ramanathan Rajaraman, Shiu-hong Kao, Angela Yao

Organizations: National University of Singapore · University of Science and Technology of China

Abstract

Spatial reasoning benchmarks evaluate vision-language models across diverse tasks, but task-level scores do not reveal which underlying capabilities account for success or failure. Each task requires recovering spatial evidence, representing geometry, and reasoning over it. We disentangle these capabilities by comparing predicted and ground-truth spatial context under a shared schema and coordinate contract. This comparison reveals four recurring sources of error: inaccurate perception, missing information in the spatial context, selection of the wrong measurement, and errors in reference frames or in tracking position and orientation. Guided by this diagnosis, we develop CROSS, a training-free library of typed geometric operators and spatial skills that function over available evidence to support reliable video spatial reasoning. The resulting library supplies verified context to non-coding VLMs or callable skills to a SpatialClaw agent. We evaluate \methodname{} on five benchmarks. \methodname{} raises the average score from 55.9% to 60.2% on ReVSI and improves the SpatialClaw result from 62.8% to 66.3% on DSI-Bench. These gains demonstrate that explicit handling of spatial conventions can repair systematic reasoning failures without additional training.

Figures & tables

Appendix figures & tables11 assets

Supplementary material from the paper’s appendix.

Appendix

Explore similar work

CardsList
  1. Reason, Then Re-reason: Cross-view Revisiting Improves Spatial Reasoning

    Jun 10, 2026Chaofan Ma, Zhenjie Mao, Yuhuan Yang +5Spatial ReasoningEgocentric Video

  2. ConsiSpace: Learning Geometric Consistency Matters for Video Spatial Reasoning

    Jul 20, 2026Ting Huang, Zhenyu Zhang, Wenyuan Huang +2Spatial ReasoningMultimodal Large Language Models

  3. ReVSI: Rebuilding Visual Spatial Intelligence Evaluation for Accurate Assessment of VLM 3D Reasoning

    Apr 27, 2026Yiming Zhang, Jiacheng Chen, Jiaqi Tan +3Visual Question Answering BenchmarksVersion