cs.CVOct 5, 2026

SpatialChain: A Benchmark for Auditing Spatial Reasoning Faithfulness in VLMs

Authors: Rafael Teixeira Sousa, Vinícius Paulo Lopes de Oliveira, Elisa Ayumi Masasi de Oliveira, Luiza Martins de Freitas Cintra, Fernanda Bufon Färber, Igor Dias Aguiar, Julia Yasmim de Almeida Nobre, Arlindo Rodrigues Galvão Filho

Organizations: Universidade Federal de Mato Grosso (UFMT) · Universidade Federal de Goiás (UFG) · Advanced Knowledge Center for Immersive Technologies (AKCIT)

Abstract

Thinking-enabled vision-language models (VLMs) report ever-higher accuracy on spatial benchmarks, yet final-answer scores cannot reveal whether a correct prediction reflects faithful spatial reasoning or a linguistic shortcut. We introduce SpatialChain, a dataset of 28,350 training and 899 test examples pairing spatially-oriented GQA questions with scene-graph-grounded reasoning chains, retained only when the generated answer matches the symbolic ground truth, and a two-axis evaluation combining objective chain-overlap metrics with a scene-graph-aware LLM judge that scores faithfulness and completeness independently of the final answer. Applied to nine thinking-enabled VLMs, the protocol surfaces three findings invisible to standard accuracy: (i) four of nine models achieve ≥\geq79% VQA accuracy while exhibiting shortcut rates above 39%, i.e., correct answers whose reasoning the judge marks as unfaithful; (ii) chain quality significantly predicts answer correctness for seven of nine models, but the two exceptions (Claude Sonnet 4.6, InternVL3.5-8B) reveal qualitatively distinct failure modes, terse output vs. verbose-decorative reasoning, that benchmark accuracy alone conflates; (iii) SFT on SpatialChain improves Qwen3-VL-8B by +6.2 pp in-domain and reduces its shortcut rate to 22%, while a stylistic specialization effect on external benchmarks motivates replay-augmented training as mitigation. The faithfulness judge is validated against 198 human-annotated items, where judge-human agreement matches human-human agreement, and against a second judge from a different provider, which preserves the model ranking (ρρ = 0.88). Data, generation scripts, and evaluation code are released at https://github.com/spatialchain/SpatialChainBenchmark.

Figures & tables

Appendix figures & tables22 assets

Supplementary material from the paper’s appendix.

Appendix

Explore similar work

CardsList
  1. ProSR: Process-Shaped Spatial Reasoning for Reliable Chain-of-Thought in VLMs

    May 25, 2026Jiangyang Li, Cong Wan, Changjie Wu +8Spatial ReasoningReasoning Trajectory

  2. Trace, Verify, and Correct: A Training-Free Framework for Spatial Reasoning in Multimodal LLMs

    Aug 5, 2026Yang Yang, Jiawei Chen, Tairan Chen +1Spatial ReasoningMultimodal Large Language Models

  3. Why Far Looks Up: Probing Spatial Representation in Vision-Language Models

    May 28, 2026Cheolhong Min, Jaeyun Jung, Daeun Lee +5Entanglement