SpatialCORE: Confidence-Aware Grounded Spatial Reasoning in Large Vision--Language Models
Authors: Rafi Ibn Sultan, Xiangyu Zhou, Md. Sajid Alam Chowdhury, Chengyin Li, Prashant Khanduri, Marco Brocanelli, Dongxiao Zhu
Organizations: Department of Computer Science, Wayne State University · Department of Radiation Oncology, Henry Ford Health · Department of Electrical and Computer Engineering, The Ohio State University · Institute for AI and Data Science, Wayne State University
Large Vision-Language Models (LVLMs) have made remarkable progress across visual perception tasks, yet spatial reasoning remains a persistent weakness, especially for questions that require reasoning over visual space. Recent spatial-reasoning methods incorporate generated grounding, where models predict bounding boxes, masks, or other localization outputs for task-relevant objects as part of their reasoning trace. However, these approaches typically optimize final-answer correctness alone, allowing correct answers to be rewarded even when the model does not reason from confidently localized task-relevant objects. We introduce SpatialCORE (Spatially COnfident REasoning), a post-training framework that turns the model's own confidence in generated grounding into a learning signal for spatial reasoning. Its central idea is to reinforce grounding that is both accurate and confident, encouraging the model to reason from confidently localized task-relevant objects. SpatialCORE realizes this through a self-regulating spatial reward that weights each predicted bounding box's matching quality by its coordinate-token confidence. An answer gate further ties grounding optimization to final-answer correctness. SpatialCORE achieves state-of-the-art results among open-source and specialized spatial reasoning models across diverse benchmarks, and transfers effectively in zero-shot settings to unseen data distributions. The source code is available at https://github.com/rafiibnsultan/SpatialCORE.
Figures & tables
Figure 1: Correct final answers do not necessarily imply confident grounding. Although both models answer correctly, (a) the baseline LVLM produces high-entropy predicted BBoxes, with dashed candidate BBoxes spread across off-target locations. (b) Our SpatialCORE produces lower-entropy predicted BBoxes, with dashed candidate BBoxes concentrated around the selected BBoxes, and more confidently localizing the task-relevant objects. Solid BBoxes denote the predicted BBoxes in the reasoning trace; dashed BBoxes denote candidate BBoxes reflected by spatial uncertainty.
Figure 2: Overview of SpatialCORE . (a) The LVLM policy samples trajectories comprising a reasoning trace with generated grounding expressed as bounding boxes (BBoxes), followed by a final answer. (b) The self-regulating spatial reward uses predicted BBox coordinate-token uncertainty to estimate the grounding confidence, which then weights each BBox’s matching quality. Predicted BBoxes are matched to pseudo-GT BBoxes using geometric overlap, label similarity, and pseudo-GT validity. For example, black sedan ahead and black sedan have high label similarity, while higher pseudo-GT validity, such as black sedan, validity: 0.8 , gives the match more weight. (c) Spatial, format, and answer rewards are composed through an answer gate to produce trajectory-level rewards, which are used to compute group-relative advantages for policy update.
Method
Average
Dynamic Reasoning
Spatial Interaction
Complex Logic
Perspective Taking
Mani- pulation
Motion Anal.
Traffic Anal.
Loca- lization
Geospa. Strategy
Pattern Rec.
Geometric Reasoning
Ego Centric
Allo Centric
Hypo- thetical
Reference Baselines
Random Choice
24.98
24.86
26.30
25.88
23.43
27.27
21.44
24.77
22.55
24.84
25.78
Human Evaluation
92.63
94.62
96.07
91.38
95.11
92.15
89.02
85.90
98.53
94.30
90.26
Proprietary Models
GPT-4.1-mini Open (2025)
48.87
64.32
56.53
59.06
60.19
56.36
29.28
30.19
72.55
39.57
39.28
Table 1: OmniSpatial Jia et al. (2025) results across 10 spatial reasoning task categories. SpatialCORE is compared against proprietary models, general open-source LVLMs, and specialized spatial reasoning models. Proprietary models are included as reference points; the best open-source or specialized result is shown in bold . Average accuracy is weighted by category sample size.
Method
Average
Question Categories
3D Geom.
Dep. & Occu.
Orientation
Relat. Posit.
Size & Scale
Spati. Navig.
Reference Baselines
Random Choice
25.00
25.00
25.00
25.00
25.00
25.00
25.00
Human Baseline
87.57
93.70
74.13
91.58
91.51
88.89
87.76
Proprietary Models
GPT-4o-mini Hurst et al. (2024)
46.50
47.06
39.00
47.03
47.17
49.60
49.79
Table 2: SpatiaLab Wasi et al. (2026) results across 6 spatial reasoning task categories in the zero-shot setting. SpatialCORE is compared against proprietary models, general open-source LVLMs, and specialized spatial reasoning models. Proprietary models are included as reference points; the best open-source or specialized result is shown in bold . Average accuracy is weighted by category sample size.
Figure 3: Qualitative examples from OmniSpatial Jia et al. (2025) (left) and SpatiaLab Wasi et al. (2026) (right). In both cases, the baseline (Qwen3-VL-8B-Thinking Bai et al. (2025) ) produces predicted BBoxes that mislocalize the task-relevant objects, leading to incorrect or poorly grounded answers. SpatialCORE generates more accurate predicted bounding boxes and reaches the correct final answer. Bounding boxes are overlaid for visualization; full reasoning traces are in the Appendix Appendix C .
Figure 4: Mean matched BBox IoU across model-specific coordinate-token entropy quartiles for the unadapted Qwen3-VL-8B-Thinking backbone (Baseline) and SpatialCORE-8B. Lower entropy corresponds to more accurate grounding after SpatialCORE post-training, whereas the baseline shows no consistent relationship.
Figure 5: Bar plots comparing answer accuracy and predicted BBox coordinate-token uncertainty for the unadapted Qwen3-VL-8B-Thinking backbone (Baseline) and SpatialCORE-8B. Lower uncertainty indicates greater confidence in BBoxes generated during reasoning.
Variant
Avg.
Traffic Anal.
Loca- lization
Geospa. Strategy
SpatialCORE-8B (Full)
56.33
54.11
62.85
51.81
w/o vision LoRA
53.00
50.59
58.10
50.00
w/o conf. weighting
52.33
49.41
58.09
49.09
w/o answer gate
54.67
52.94
60.00
50.91
w/o pseudo-GT validity
53.33
51.76
59.05
49.09
Table 3: Ablation study on the spatial interaction subset of OmniSpatial Jia et al. (2025) . Average is weighted by category sample size. Best result is bold .
Appendix figures & tables6 assets
Supplementary material from the paper’s appendix.
Appendix
Hyperparameter
Value
Notes
Model and architecture
Base model
Qwen3-VL-8B-Thinking
Reasoning backbone
LoRA rank
32
Language-side LoRA
LoRA alpha
64
Language-side LoRA
LoRA dropout
0.05
–
LoRA target modules
Default
Architecture defaults
Appendix
Table 4: Hyperparameter settings for post-training SpatialCORE-8B with GRPO, including model configuration, training setup, optimization parameters, and reward weights.
Figure 6: Manual audit precision of pseudo-GT BBoxes across Grounding DINO confidence ranges. Points show precision and whiskers show 95% confidence intervals; labels give audited BBox counts. The dashed line marks overall precision.
Training setting
Acc. (%)
Base
43.90
SpatialCORE-8B ( 40% corrupted)
46.70
SpatialCORE-8B (original)
48.46
Appendix
Table 5: Robustness to pseudo-GT corruption on the full OmniSpatial test set. Base denotes Qwen3-VL-8B-Thinking.
Figure 7: System prompt used to standardize generated grounding trajectories during SpatialCORE training. Continued on the next page.
Figure 8: System prompt used to standardize generated grounding trajectories during SpatialCORE training, continued.
Figure 9: Qualitative example of confidence-aware grounded spatial reasoning with SpatialCORE. The model first generates predicted BBoxes for the task-relevant objects, including thermostat , scissor , lamp , and cup , and then reasons over their localized positions relative to the chair. By producing confident generated grounding during the reasoning trace, SpatialCORE identifies the wall-mounted thermostat as the hardest object to reach and selects the correct final answer. Bounding boxes are overlaid only for visualization.
Department of Computer Science, Wayne State University · Department of Radiation Oncology, Henry Ford Health · Department of Electrical and Computer Engineering, The Ohio State University +1