SpatialCORE: Confidence-Aware Grounded Spatial Reasoning in Large Vision--Language Models
Authors: Rafi Ibn Sultan, Xiangyu Zhou, Md. Sajid Alam Chowdhury, Chengyin Li, Prashant Khanduri, Marco Brocanelli, Dongxiao Zhu
Organizations: Department of Computer Science, Wayne State University · Department of Radiation Oncology, Henry Ford Health · Department of Electrical and Computer Engineering, The Ohio State University · Institute for AI and Data Science, Wayne State University
Large Vision-Language Models (LVLMs) have made remarkable progress across visual perception tasks, yet spatial reasoning remains a persistent weakness, especially for questions that require reasoning over visual space. Recent spatial-reasoning methods incorporate generated grounding, where models predict bounding boxes, masks, or other localization outputs for task-relevant objects as part of their reasoning trace. However, these approaches typically optimize final-answer correctness alone, allowing correct answers to be rewarded even when the model does not reason from confidently localized task-relevant objects. We introduce SpatialCORE (Spatially COnfident REasoning), a post-training framework that turns the model's own confidence in generated grounding into a learning signal for spatial reasoning. Its central idea is to reinforce grounding that is both accurate and confident, encouraging the model to reason from confidently localized task-relevant objects. SpatialCORE realizes this through a self-regulating spatial reward that weights each predicted bounding box's matching quality by its coordinate-token confidence. An answer gate further ties grounding optimization to final-answer correctness. SpatialCORE achieves state-of-the-art results among open-source and specialized spatial reasoning models across diverse benchmarks, and transfers effectively in zero-shot settings to unseen data distributions. The source code is available at https://github.com/rafiibnsultan/SpatialCORE.
Figures & tables
Figure 1: Correct final answers do not necessarily imply confident grounding. Although both models answer correctly, (a) the baseline LVLM produces high-entropy predicted BBoxes, with dashed candidate BBoxes spread across off-target locations. (b) Our SpatialCORE produces lower-entropy predicted BBoxes, with dashed candidate BBoxes concentrated around the selected BBoxes, and more confidently localizing the task-relevant objects. Solid BBoxes denote the predicted BBoxes in the reasoning trace; dashed BBoxes denote candidate BBoxes reflected by spatial uncertainty.
Figure 2: Overview of SpatialCORE . (a) The LVLM policy samples trajectories comprising a reasoning trace with generated grounding expressed as bounding boxes (BBoxes), followed by a final answer. (b) The self-regulating spatial reward uses predicted BBox coordinate-token uncertainty to estimate the grounding confidence, which then weights each BBox’s matching quality. Predicted BBoxes are matched to pseudo-GT BBoxes using geometric overlap, label similarity, and pseudo-GT validity. For example, black sedan ahead and black sedan have high label similarity, while higher pseudo-GT validity, such as black sedan, validity: 0.8 , gives the match more weight. (c) Spatial, format, and answer rewards are composed through an answer gate to produce trajectory-level rewards, which are used to compute group-relative advantages for policy update.
Method
Average
Dynamic Reasoning
Spatial Interaction
Complex Logic
Perspective Taking
Mani- pulation
Motion Anal.
Traffic Anal.
Loca- lization
Geospa. Strategy
Pattern Rec.
Geometric Reasoning
Ego Centric
Allo Centric
Hypo- thetical
Reference Baselines
Random Choice
24.98
24.86
26.30
25.88
23.43
27.27
21.44
24.77
22.55
24.84
25.78
Human Evaluation
92.63
94.62
96.07
91.38
95.11
92.15
89.02
85.90
98.53
94.30
90.26
Proprietary Models
GPT-4.1-mini Open (2025)
48.87
64.32
56.53
59.06
60.19
56.36
29.28
30.19
72.55
39.57
39.28
Table 1: OmniSpatial Jia et al. (2025) results across 10 spatial reasoning task categories. SpatialCORE is compared against proprietary models, general open-source LVLMs, and specialized spatial reasoning models. Proprietary models are included as reference points; the best open-source or specialized result is shown in bold . Average accuracy is weighted by category sample size.
Method
Average
Question Categories
3D Geom.
Dep. & Occu.
Orientation
Relat. Posit.
Size & Scale
Spati. Navig.
Reference Baselines
Random Choice
25.00
25.00
25.00
25.00
25.00
25.00
25.00
Human Baseline
87.57
93.70
74.13
91.58
91.51
88.89
87.76
Proprietary Models
GPT-4o-mini Hurst et al. (2024)
46.50
47.06
39.00
47.03
47.17
49.60
49.79
Table 2: SpatiaLab Wasi et al. (2026) results across 6 spatial reasoning task categories in the zero-shot setting. SpatialCORE is compared against proprietary models, general open-source LVLMs, and specialized spatial reasoning models. Proprietary models are included as reference points; the best open-source or specialized result is shown in bold . Average accuracy is weighted by category sample size.
Figure 3: Qualitative examples from OmniSpatial Jia et al. (2025) (left) and SpatiaLab Wasi et al. (2026) (right). In both cases, the baseline (Qwen3-VL-8B-Thinking Bai et al. (2025) ) produces predicted BBoxes that mislocalize the task-relevant objects, leading to incorrect or poorly grounded answers. SpatialCORE generates more accurate predicted bounding boxes and reaches the correct final answer. Bounding boxes are overlaid for visualization; full reasoning traces are in the Appendix Appendix C .
Figure 4: Mean matched BBox IoU across model-specific coordinate-token entropy quartiles for the unadapted Qwen3-VL-8B-Thinking backbone (Baseline) and SpatialCORE-8B. Lower entropy corresponds to more accurate grounding after SpatialCORE post-training, whereas the baseline shows no consistent relationship.
Figure 5: Bar plots comparing answer accuracy and predicted BBox coordinate-token uncertainty for the unadapted Qwen3-VL-8B-Thinking backbone (Baseline) and SpatialCORE-8B. Lower uncertainty indicates greater confidence in BBoxes generated during reasoning.
Variant
Avg.
Traffic Anal.
Loca- lization
Geospa. Strategy
SpatialCORE-8B (Full)
56.33
54.11
62.85
51.81
w/o vision LoRA
53.00
50.59
58.10
50.00
w/o conf. weighting
52.33
49.41
58.09
49.09
w/o answer gate
54.67
52.94
60.00
50.91
w/o pseudo-GT validity
53.33
51.76
59.05
49.09
Table 3: Ablation study on the spatial interaction subset of OmniSpatial Jia et al. (2025) . Average is weighted by category sample size. Best result is bold .
Appendix figures & tables6 assets
Supplementary material from the paper’s appendix.
Appendix
Hyperparameter
Value
Notes
Model and architecture
Base model
Qwen3-VL-8B-Thinking
Reasoning backbone
LoRA rank
32
Language-side LoRA
LoRA alpha
64
Language-side LoRA
LoRA dropout
0.05
–
LoRA target modules
Default
Architecture defaults
Appendix
Table 4: Hyperparameter settings for post-training SpatialCORE-8B with GRPO, including model configuration, training setup, optimization parameters, and reward weights.
Figure 6: Manual audit precision of pseudo-GT BBoxes across Grounding DINO confidence ranges. Points show precision and whiskers show 95% confidence intervals; labels give audited BBox counts. The dashed line marks overall precision.
Training setting
Acc. (%)
Base
43.90
SpatialCORE-8B ( 40% corrupted)
46.70
SpatialCORE-8B (original)
48.46
Appendix
Table 5: Robustness to pseudo-GT corruption on the full OmniSpatial test set. Base denotes Qwen3-VL-8B-Thinking.
Figure 7: System prompt used to standardize generated grounding trajectories during SpatialCORE training. Continued on the next page.
Figure 8: System prompt used to standardize generated grounding trajectories during SpatialCORE training, continued.
Figure 9: Qualitative example of confidence-aware grounded spatial reasoning with SpatialCORE. The model first generates predicted BBoxes for the task-relevant objects, including thermostat , scissor , lamp , and cup , and then reasons over their localized positions relative to the chair. By producing confident generated grounding during the reasoning trace, SpatialCORE identifies the wall-mounted thermostat as the hardest object to reach and selects the correct final answer. Bounding boxes are overlaid only for visualization.
Vision-Language Models (VLMs) have demonstrated remarkable capabilities in multimodal tasks, yet they still exhibit poor ability in spatial reasoning. Existing training-dependent and training-free enhancement methods suffer from high computational costs with catastrophic forgetting and internal mechanism interference that compromises general capabilities, respectively. In this work, we first verify two key hypotheses: appropriate geometric image transformation and query-reversal transformation can recover incorrect spatial predictions, and correct predictions exhibit higher relation-token confidence than incorrect ones. Based on these findings, we propose INTCORT, a training-free spatial reasoning enhancement framework that constructs multiple inference views through input transformations and aggregates their predictions via relation-token confidence routing, without modifying the VLM's internal mechanisms. Experimental results on several commonly-used benchmarks demonstrate that INTCORT substantially improves spatial reasoning accuracy across diverse VLMs, achieving an average improvement of 10.01% over all models and benchmarks. Compared with prior works, INTCORT achieves superior performance with improvements of up to 25.01%.
Haoran Sun, Jingqi Xu, Yanhui Li +3
The University of Hong Kong · University of Southern California · China Telecom +3
Multimodal large language models (MLLMs) have achieved remarkable progress in vision-language tasks, but continue to struggle with spatial reasoning. Existing spatial MLLMs rely on large-scale datasets, explicit 3D inputs, architecture-specific modifications, or sparse Reinforcement Learning (RL) methods that provide insufficient guidance for spatially-grounded reasoning. We introduce SpatialThinker. To our knowledge, it is the first MLLM unifying Scene Graph Generation (SGG) and visual reasoning in a single pass via online RL. The model simulates human-like spatial perception by constructing a mental scene graph of task-relevant objects and relations, and reasoning toward an answer via dense spatial rewards. Our contributions are threefold: (1) SGG-grounded reasoning: integrating SGG directly within the reasoning chain rather than as a disjoint preprocessing step; (2) STVQA-7K: a high-quality spatial VQA training dataset via a scalable synthesis pipeline; and (3) a dense spatial reward design that enforces structured grounding during RL and generalizes to improve broad visual perception. SpatialThinker-7B achieves 3.6× larger gains over SFT and 1.7× better in- and out-of-distribution generalization than sparse RL. Trained on only 7K samples, SpatialThinker-7B matches GPT-5 and outperforms GPT-4o, while SpatialThinker-30B surpasses both GPT-5 and Claude 4 Sonnet on average across 14 spatial and real-world benchmarks, demonstrating that structured spatial grounding with reward-aligned reasoning enables robust spatial understanding with limited data.
Hunar Batra, Haoqin Tu, Hardy Chen +3
University of Oxford · University of California, Santa Cruz
Large Vision-Language Models (LVLMs) commonly perform spatial reasoning through chain-of-thought (CoT), encoding intermediate reasoning as autoregressive sequences of discrete language tokens. Such hard thinking requires committing to a single token at each step, even when the correct spatial interpretation remains uncertain. This early commitment constitutes premature discretization: an incorrect token selection can propagate errors through subsequent reasoning. We propose Soft Spatial Reasoning, a post-training framework that introduces soft thinking for spatial tasks in LVLMs. At each intermediate reasoning step, the LVLM forms a continuous soft state by mixing token embeddings rather than selecting a single token, allowing multiple candidate continuations to influence the next step. The appropriate degree of softness, however, can vary across reasoning steps: retaining multiple candidates may preserve a useful spatial interpretation, but if those candidates imply conflicting spatial relations, mixing them may interfere with subsequent reasoning. At the core of Soft Spatial Reasoning is AdaptSoft, a controller that uses the current hidden state and predictive uncertainty to adapt the degree of softness at each reasoning step. To train AdaptSoft, we introduce a gradient-alignment learning objective that provides a step-specific learning signal for softness control without intermediate reasoning supervision. Across diverse spatial benchmarks, Soft Spatial Reasoning outperforms hard and fixed-soft CoT baselines using the same backbone, as well as a range of existing LVLMs. The source code is available at https://github.com/rafiibnsultan/Soft_Spatial_Reasoning
Rafi Ibn Sultan, Md. Sajid Alam Chowdhury, Saleh Zare Zade +4
Department of Computer Science, Wayne State University · Department of Radiation Oncology, Henry Ford Health · Department of Electrical and Computer Engineering, The Ohio State University +1