SpatialChain: A Benchmark for Auditing Spatial Reasoning Faithfulness in VLMs
Organizations: Universidade Federal de Mato Grosso (UFMT) · Universidade Federal de Goiás (UFG) · Advanced Knowledge Center for Immersive Technologies (AKCIT)
Abstract
Thinking-enabled vision-language models (VLMs) report ever-higher accuracy on spatial benchmarks, yet final-answer scores cannot reveal whether a correct prediction reflects faithful spatial reasoning or a linguistic shortcut. We introduce SpatialChain, a dataset of 28,350 training and 899 test examples pairing spatially-oriented GQA questions with scene-graph-grounded reasoning chains, retained only when the generated answer matches the symbolic ground truth, and a two-axis evaluation combining objective chain-overlap metrics with a scene-graph-aware LLM judge that scores faithfulness and completeness independently of the final answer. Applied to nine thinking-enabled VLMs, the protocol surfaces three findings invisible to standard accuracy: (i) four of nine models achieve 79% VQA accuracy while exhibiting shortcut rates above 39%, i.e., correct answers whose reasoning the judge marks as unfaithful; (ii) chain quality significantly predicts answer correctness for seven of nine models, but the two exceptions (Claude Sonnet 4.6, InternVL3.5-8B) reveal qualitatively distinct failure modes, terse output vs. verbose-decorative reasoning, that benchmark accuracy alone conflates; (iii) SFT on SpatialChain improves Qwen3-VL-8B by +6.2 pp in-domain and reduces its shortcut rate to 22%, while a stylistic specialization effect on external benchmarks motivates replay-augmented training as mitigation. The faithfulness judge is validated against 198 human-annotated items, where judge-human agreement matches human-human agreement, and against a second judge from a different provider, which preserves the model ranking ( = 0.88). Data, generation scripts, and evaluation code are released at https://github.com/spatialchain/SpatialChainBenchmark.
Figures & tables
| Split | Source | Samples | Images | Yes / No / Attr. | Yes-rate | Avg. chain (w) |
| Train | GQA train | 28,350 | 19,263 | 16,516 / 9,926 / 1,908 | 0.58 | 115 |
| Test | GQA val | 899 | 719 | 566 / 329 / 4 | 0.63 | 145 |
| Model | Acc | F1 | Yes | No | R-1 | R-L | F1 tok | Len |
| Qwen3-VL-8B FT | 85.03 | 88.68 | 92.76 | 71.73 | 0.654 | 0.443 | 0.643 | 0.90 |
| Qwen3-VL-4B FT | 82.23 | 86.67 | 91.34 | 66.57 | 0.657 | 0.450 | 0.646 | 0.89 |
| GLM-4.1-9B | 81.34 | 84.49 | 80.39 | 82.98 | 0.438 | 0.301 | 0.421 | 1.47 |
| InternVL3.5-4B † | 81.34 | 87.43 | 87.28 | 71.12 | 0.314 | 0.227 | 0.305 | 3.01 |
| Claude Sonnet 4.6 † | 79.89 | 85.15 | 87.63 | 66.57 | 0.359 | 0.258 | 0.354 | 0.36 |
| InternVL3.5-8B † | 79.33 | 85.29 | 86.57 | 66.87 | 0.323 | 0.228 | 0.316 | 7.07 |
| Model | R-1 ✓ | R-1 × | R-1 | F1 ✓ | F1 × | |
| Qwen3-VL-4B | 0.418 | 0.351 | +0.067 | 0.406 | 0.342 | |
| Qwen3-VL-8B | 0.439 | 0.373 | +0.066 | 0.426 | 0.362 | |
| GLM-4.1-9B | 0.447 | 0.399 | +0.048 | 0.429 | 0.385 | |
| Qwen3-VL-4B FT | 0.664 | 0.625 | +0.039 | 0.653 | 0.612 | |
| Qwen3-VL-8B FT | 0.660 | 0.626 | +0.034 | 0.649 | 0.612 | |
| InternVL3.5-4B | 0.320 | 0.288 | +0.031 | 0.311 | 0.282 |
| Model | Acc | Faith. | Compl. | Pass | Shortcut |
| Qwen3-VL-8B FT | 85.03 | 0.630 | 0.700 | 0.820 | 0.222 |
| Qwen3-VL-4B FT | 82.23 | 0.631 | 0.708 | 0.802 | 0.194 |
| GLM-4.1-9B | 81.34 | 0.527 | 0.583 | 0.809 | 0.453 |
| Claude Sonnet 4.6 | 79.89 | 0.541 | 0.582 | 0.794 | 0.390 |
| Qwen3-VL-4B | 78.44 | 0.585 | 0.658 | 0.776 | 0.264 |
| Qwen3-VL-8B | 78.88 | 0.571 | 0.658 | 0.765 | 0.260 |
Appendix figures & tables22 assets
Supplementary material from the paper’s appendix.
Appendix
| Group | GQA types |
| Spatial relation verification | relVerify , relVerifyCop , relVerifyCr |
| Existential with relation | existRelS , existRelSC , existRelSRC |
| Place verification | placeVerify , placeVerifyC |
| Two-object comparison | twoDifferent , twoDifferentC , twoSame , twoSameC |
| Category comparison | diffAnimals , diffAnimalsC , diffGender |
| Attribute / ordinal selection | chooseAttr , compare |
| Relation | Count | % | Yes-rate |
| left_of | 9,832 | 34.7 | 0.57 |
| right_of | 7,919 | 27.9 | 0.58 |
| in_front_of | 1,987 | 7.0 | 0.54 |
| behind | 1,406 | 5.0 | 0.55 |
| near | 347 | 1.2 | 0.61 |
| above | 763 | 2.7 | 0.71 |
| Hyperparameter | Value |
| Base model | Qwen3-VL-8B-Thinking |
| Quantization | 4-bit NF4 (BitsAndBytes) |
| Quantization compute dtype | bfloat16 |
| Double quantization | ✓ |
| Attention implementation | Flash Attention 2 |
| LoRA rank | 16 |
| Hyperparameter | Value |
| Training examples | 28,350 |
| Epochs | 2 |
| Per-device batch size | 6 |
| Gradient accumulation steps | 2 |
| Effective batch size | 12 |
| Learning rate |
| Model | Acc | R-1 | R-2 | R-L | F1 tok | Prec tok | Rec tok | BLEU-1 | Entity-ov. | Step-mk. |
| Qwen3-VL-4B FT | 82.23 | 0.657 | 0.407 | 0.450 | 0.646 | 0.701 | 0.614 | 0.570 | 0.481 | 0.26 |
| Qwen3-VL-8B FT | 85.03 | 0.654 | 0.402 | 0.443 | 0.643 | 0.694 | 0.615 | 0.568 | 0.484 | 0.45 |
| GLM-4.1-9B | 81.34 | 0.438 | 0.212 | 0.301 | 0.421 | 0.474 | 0.470 | 0.302 | 0.305 | 1.34 |
| Qwen3-VL-8B | 78.88 | 0.425 | 0.218 | 0.293 | 0.412 | 0.353 | 0.597 | 0.328 | 0.368 | 2.15 |
| Qwen3-VL-4B | 78.44 | 0.403 | 0.208 | 0.279 | 0.392 | 0.323 | 0.613 | 0.303 | 0.379 | 2.77 |
| Claude 4.6 | 79.89 | 0.359 | 0.187 | 0.258 | 0.354 | 0.683 | 0.248 | 0.127 | 0.253 | 0.00 |
| Model | left_of | right_of | in_front_of | behind | above | below | next_to | other |
| Qwen3-VL-8B FT | 0.836 | 0.868 | 0.800 | 1.000 | 0.833 | 0.500 | 0.750 | 0.923 |
| Qwen3-VL-4B FT | 0.816 | 0.825 | 0.800 | 0.667 | 0.833 | 1.000 | 1.000 | 0.923 |
| GLM-4.1-9B | 0.812 | 0.808 | 0.933 | 0.667 | 0.833 | 0.500 | 0.750 | 1.000 |
| InternVL3.5-4B | 0.803 | 0.815 | 1.000 | 0.667 | 0.667 | 1.000 | 0.750 | 1.000 |
| Qwen3-VL-4B | 0.796 | 0.778 | 0.733 | 0.667 | 0.833 | 0.000 | 0.750 | 0.769 |
| Qwen3-VL-8B | 0.789 | 0.788 | 0.667 | 1.000 | 0.833 | 0.500 | 0.750 | 0.923 |
| Checkpoint | Traj. | Action | Spatial | State | Task | Multi-v. | Point. | Other | Overall |
| Base | 0.409 | 0.431 | 0.524 | 0.473 | 0.500 | 0.405 | 0.382 | 0.429 | 0.453 |
| Step 63 | 0.288 | 0.375 | 0.488 | 0.527 | 0.368 | 0.297 | 0.353 | 0.143 | 0.388 |
| Step 313 | 0.424 | 0.458 | 0.464 | 0.418 | 0.500 | 0.324 | 0.441 | 0.071 | 0.425 |
| Step 625 | 0.167 | 0.417 | 0.464 | 0.455 | 0.474 | 0.324 | 0.412 | 0.214 | 0.380 |
| Step 1875 | 0.303 | 0.417 | 0.405 | 0.436 | 0.421 | 0.351 | 0.324 | 0.429 | 0.385 |
| Step 2813 | 0.303 | 0.389 | 0.405 | 0.436 | 0.342 | 0.351 | 0.441 | 0.143 | 0.372 |
| Checkpoint | Left | Right | Above | Under | Close | Far | Overall |
| Base | 0.943 | 0.905 | 0.782 | 0.772 | 0.699 | 0.640 | 0.791 |
| Step 63 | 0.937 | 0.913 | 0.797 | 0.698 | 0.676 | 0.608 | 0.773 |
| Step 313 | 0.929 | 0.873 | 0.819 | 0.701 | 0.670 | 0.582 | 0.763 |
| Step 625 | 0.904 | 0.845 | 0.809 | 0.724 | 0.642 | 0.609 | 0.757 |
| Step 1875 | 0.904 | 0.858 | 0.809 | 0.704 | 0.613 | 0.567 | 0.744 |
| Step 2813 | 0.909 | 0.873 | 0.807 | 0.721 | 0.624 | 0.601 | 0.757 |
| Model | Full prediction | Correct |
|---|---|---|
| Reference (scene graph) | “Looking at the image, let me trace through this step-by-step: (1) Locating the knife: I can see a knife on the left side of the plate, positioned near the bottom left. (2) Finding the bread to the right of the knife: there is a large piece of bread that sits on the plate to the right of where the knife is placed. (3) Examining what is to the right of that bread: to the right of the bread I can see what appears to be gray birds standing on the plate. These are birds with wings and legs visible. (4) Looking for kittens: I do not see any kittens anywhere in the image. The animals present in the scene are birds — specifically gray colored birds that appear to be standing on or near the bread on the plate.” Answer: No, there is a bird to the right of the bread. | — |
| Qwen3-VL-4B FT (fine-tuned) | “Looking at the image, I need to trace through this step by step: (1) First, I can see a knife on the left side of the image, positioned near the edge. (2) To the right of the knife, there is bread visible in the center-left portion of the image. (3) To the right of that bread, I can see a white plate with a bird perched on it. (4) Looking further to the right of the bread (and the plate with the bird), I see what appears to be a small bird on the table, but examining the entire right side of the bread, I do not see any kittens anywhere in the image. The image shows a table setting with a woman, food items including bread, a plate, a bird, and other dining elements. There are no kittens visible anywhere in the scene, let alone to the right of the bread.” Answer: No, there is a bird to the right of the bread. | ✓ |
| Qwen3-VL-8B FT (fine-tuned) | “Looking at the image, I can see a white plate with various food items on it. On the left side, there’s a knife, and to the right of it is a piece of bread. Moving further to the right of the bread, I can see what appears to be a small brown bird standing on the plate. However, I don’t see any kittens anywhere in this image. The only animal visible to the right of the bread is a small brown bird. There are no kittens present in the scene.” Answer: No, there are no kittens visible in the image. | ✓ |
| Qwen3-VL-4B (base) | “So, let us break this down. First, the question is asking if there are kittens to the right of the bread that is to the right of the knife. Let us identify each element. First, the knife: there is a knife on the left side of the plate, near the bread. The bread is the sandwich in the center of the plate. Now, ‘the bread that is to the right of the knife’ — so the knife is on the left, the bread is to the right of that knife. Then, we need to look to the right of that bread. What is there? On the plate, to the right of the bread (sandwich), there is a sauce (red sauce, maybe tomato) and some other food. But the question is about kittens. Are there any kittens in the image? Let us check the image. The image shows a plate with a sandwich, some sauce, a knife, and two small birds (sparrows) on the plate or near it. There are no kittens visible. The birds are sparrows, not kittens. So, to the right of the bread (which is the sandwich), there are no kittens. The birds are there, but they are birds, not kittens. So the answer is no.” Answer: No. | ✓ |
| Qwen3-VL-8B (base) | “Got it, let’s break this down. First, identify the elements. The question is about kittens to the right of the bread that’s to the right of the knife. Let’s look at the image. First, find the knife. There’s a knife on the left side of the plate, next to the fork. Then, the bread: the hamburger bun is on the plate, to the right of the knife. Now, to the right of that bread (hamburger bun), what’s there? There’s a bird, not a kitten. The image shows birds, not kittens. So no kittens are present. The question asks if there are kittens to the right of the bread (hamburger) which is right of the knife. Since there are birds, not kittens, the answer is no.” Answer: The image shows birds (not kittens) to the right of the hamburger bun, which is to the right of the knife. No kittens are visible. Answer: No, there are no kittens in the image | ✓ |
| GLM-4.1-9B | “Got it, let’s analyze the image. First, look for the knife, then the bread to its right, then check if there are kittens to the right of that bread. The image shows a plate with food, a knife on the left, bread (the bun) next to it, and then birds, not kittens. So there are no kittens in that area. Therefore, the answer would be no.” Answer: To determine this, first locate the knife on the left side of the plate, then identify the bread (bun) to the right of the knife. Next, check for kittens to the right of that bread. The area to the right of the bread shows birds, not kittens. Therefore, the answer is No. | ✓ |
| Dimension | human human | judge individual human | judge human consensus |
| Answer correctness | 0.673 | 0.655 | 0.679 |
| Reasoning faithfulness | 0.443 | 0.488 | 0.574 |
| Reasoning completeness | 0.416 | 0.551 | 0.643 |
| Model | Claude Sonnet 4.6 judge | GPT-5.6 judge |
| Qwen3-VL-8B FT | 0.222 | 0.466 |
| Qwen3-VL-4B FT | 0.194 | 0.396 |
| GLM-4.1-9B | 0.453 | 0.595 |
| Claude Sonnet 4.6 | 0.390 | 0.546 |
| Qwen3-VL-4B | 0.264 | 0.574 |
| Qwen3-VL-8B | 0.260 | 0.588 |
| Model | |||||
| Qwen3-VL-8B FT | 0.026 | 0.029 | 0.222 | 0.227 | 0.358 |
| Qwen3-VL-4B FT | 0.019 | 0.020 | 0.194 | 0.195 | 0.305 |
| GLM-4.1-9B | 0.003 | 0.003 | 0.453 | 0.455 | 0.627 |
| Claude Sonnet 4.6 | 0.008 | 0.009 | 0.390 | 0.390 | 0.473 |
| Qwen3-VL-4B | 0.010 | 0.010 | 0.264 | 0.265 | 0.433 |
| Qwen3-VL-8B | 0.018 | 0.018 | 0.260 | 0.265 | 0.427 |
| Subset | Pearson | Spearman | Kendall | |
| All models | 9 | ( ) | ( ) | ( ) |
| Zero-shot only | 7 | ( ) | ( ) | ( ) |
| Qwen3 family only | 4 | ( ) | ( ) | ( ) |
| Model | Acc single | Acc multi | Acc | RF single | RF multi | RF |
| Qwen3-VL-8B FT | 0.847 | 0.845 | 0.630 | 0.630 | ||
| Qwen3-VL-4B FT | 0.834 | 0.809 | 0.641 | 0.623 | ||
| GLM-4.1-9B | 0.810 | 0.815 | 0.517 | 0.533 | ||
| Claude Sonnet 4.6 | 0.777 | 0.823 | 0.529 | 0.550 | ||
| Qwen3-VL-4B | 0.764 | 0.805 | 0.566 | 0.597 | ||
| Qwen3-VL-8B | 0.748 | 0.815 | 0.550 | 0.586 |