Authors: Rafi Ibn Sultan, Md. Sajid Alam Chowdhury, Saleh Zare Zade, Chengyin Li, Prashant Khanduri, Marco Brocanelli, Dongxiao Zhu
Organizations: Department of Computer Science, Wayne State University · Department of Radiation Oncology, Henry Ford Health · Department of Electrical and Computer Engineering, The Ohio State University · Institute for AI and Data Science, Wayne State University
Large Vision-Language Models (LVLMs) commonly perform spatial reasoning through chain-of-thought (CoT), encoding intermediate reasoning as autoregressive sequences of discrete language tokens. Such hard thinking requires committing to a single token at each step, even when the correct spatial interpretation remains uncertain. This early commitment constitutes premature discretization: an incorrect token selection can propagate errors through subsequent reasoning. We propose Soft Spatial Reasoning, a post-training framework that introduces soft thinking for spatial tasks in LVLMs. At each intermediate reasoning step, the LVLM forms a continuous soft state by mixing token embeddings rather than selecting a single token, allowing multiple candidate continuations to influence the next step. The appropriate degree of softness, however, can vary across reasoning steps: retaining multiple candidates may preserve a useful spatial interpretation, but if those candidates imply conflicting spatial relations, mixing them may interfere with subsequent reasoning. At the core of Soft Spatial Reasoning is AdaptSoft, a controller that uses the current hidden state and predictive uncertainty to adapt the degree of softness at each reasoning step. To train AdaptSoft, we introduce a gradient-alignment learning objective that provides a step-specific learning signal for softness control without intermediate reasoning supervision. Across diverse spatial benchmarks, Soft Spatial Reasoning outperforms hard and fixed-soft CoT baselines using the same backbone, as well as a range of existing LVLMs. The source code is available at https://github.com/rafiibnsultan/Soft_Spatial_Reasoning
Figures & tables
Figure 1: Hard versus adaptive soft thinking for spatial reasoning. (a) Hard thinking commits to one token at each CoT step, allowing early errors to propagate. (b) Our Soft Spatial Reasoning forms continuous intermediate states by mixing token embeddings, carrying information from multiple candidate continuations through the CoT. AdaptSoft adjusts the mixture’s softness at each step based on the current reasoning state and predictive uncertainty.
Figure 2: Overview of Soft Spatial Reasoning . (a) The LVLM policy is post-trained to carry multiple candidate continuations through soft states during reasoning, then generate the final answer as discrete tokens. Rollout rewards yield group-relative advantages for the GRPO update. (b) At each reasoning step, AdaptSoft sets the mixture temperature from the current hidden state and candidate-distribution entropy. The resulting mixture weights combine token embeddings into the next soft state.
Figure 3: Gradient-Alignment Learning. At each soft reasoning step t , the temperature τi,t controls the state si,t , whose contribution gi,t to the LVLM output-layer gradient is compared with a reference gradient Gref computed from a disjoint subset of rollouts. The resulting alignment scores αi,t are aggregated to form the AdaptSoft loss.
Method
Average
Dynamic Reasoning
Spatial Interaction
Complex Logic
Perspective Taking
Mani- pulation
Motion Anal.
Traffic Anal.
Loca- lization
Geospa. Strategy
Pattern Rec.
Geometric Reasoning
Ego Centric
Allo Centric
Hypo- thetical
Reference Baselines
Random Choice
24.98
24.86
26.30
25.88
23.43
27.27
21.44
24.77
22.55
24.84
25.78
Human Evaluation
92.63
94.62
96.07
91.38
95.11
92.15
89.02
85.90
98.53
94.30
90.26
Proprietary Models
GPT-4.1-mini Open (2025)
48.87
64.32
56.53
59.06
60.19
56.36
29.28
30.19
72.55
39.57
39.28
Table 1: OmniSpatial Jia et al. (2025) results across 10 spatial reasoning task categories. Soft Spatial Reasoning is compared against proprietary models, general open-source LVLMs, soft thinking models, and specialized spatial reasoning models. Proprietary models are included as reference points; the best open-source or specialized result is shown in bold . Average accuracy is weighted by category sample size.
Figure 4: Soft Spatial Reasoning on a randomly selected example from OmniSpatial. AdaptSoft varies softness across the CoT. For readability, greedy decoding is used only to display the soft states as text. Shading shows each span’s mean temperature ( blue : lower; orange : higher). The plot below shows temperature at every reasoning step and labels several words with their temperatures.
Method
Average
Question Categories
3D Geom.
Dep. & Occu.
Orientation
Relat. Posit.
Size & Scale
Spati. Navig.
Reference Baselines
Random Choice
25.00
25.00
25.00
25.00
25.00
25.00
25.00
Human Baseline
87.57
93.70
74.13
91.58
91.51
88.89
87.76
Proprietary Models
GPT-4o-mini Hurst et al. (2024)
46.50
47.06
39.00
47.03
47.17
49.60
49.79
Table 2: SpatiaLab Wasi et al. (2026) results across 6 spatial reasoning task categories in the zero-shot setting. Proprietary models are included as reference points; the best open-source or specialized result is shown in bold . Average accuracy is weighted by category sample size.
Method
Overall
Rotation
Among
Around
Reference Baseline
Random Choice
32.35
36.36
32.29
30.66
Proprietary Models
Gemini-2.5-Pro Team et al. (2023)
47.05
85.50
25.95
38.40
Claude-4-Sonnet Anthropic (2024)
44.75
48.42
44.21
47.62
Open-weight Models
Table 3: Zero-shot results on MindCube Wang et al. (2026a) across three spatial mental-modeling settings. Proprietary models are included as reference points, and Overall is weighted by setting size.
Variant
Avg.
Ego centric
Allo centric
Hypo thetical
Soft Spatial Reasoning (Full)
42.75
81.37
32.53
41.46
w/o hi,t in AdaptSoft
39.22
77.45
29.52
36.14
w/o predictive uncertainty
40.64
78.43
30.85
38.55
w/o gradient alignment
38.68
76.47
29.26
34.94
w/o gradient centering
41.18
79.41
31.12
39.76
Table 4: OmniSpatial Perspective Taking ablations Jia et al. (2025) . All variants share the backbone and policy GRPO setup, with post-training and evaluation on the corresponding training and test splits. Averages are sample-weighted.
Appendix figures & tables5 assets
Supplementary material from the paper’s appendix.
Appendix
Figure 5: AdaptSoft Temperatures During Training. The line shows the mean temperature τi,t across soft steps in each rollout batch; shading spans their minimum and maximum. The dashed line marks τ0=0.5 .
Figure 6: AdaptSoft Temperature and Predictive Uncertainty. Temperatures and predictive uncertainty are measured across soft reasoning steps during inference on the test set. Color shows the number of steps on a logarithmic scale. The solid line shows median temperature in each entropy bin; dashed lines mark the 10th and 90th percentiles.
Figure 7: AdaptSoft Varies Softness Within CoTs. The histogram shows the 90th–10th percentile temperature spread within each test-set reasoning trace generated during inference. The solid line marks the median within-trace spread; the dashed line marks the corresponding spread across trace-median temperatures.
Hyperparameter
Value
Notes
Model configuration
Backbone
Qwen3-VL-8B-Thinking
–
Language-model adaptation
Full fine-tuning
No LoRA
Numerical precision
BF16
FSDP2 mixed precision
Parameter sharding
FSDP2
Parameter and optimizer offload
Data preparation
Appendix
Table 5: Post-training and evaluation configuration for Soft Spatial Reasoning.
Figure 8: System prompt used to standardize the reasoning format during Soft Spatial Reasoning training.
Large Vision-Language Models (LVLMs) have made remarkable progress across visual perception tasks, yet spatial reasoning remains a persistent weakness, especially for questions that require reasoning over visual space. Recent spatial-reasoning methods incorporate generated grounding, where models predict bounding boxes, masks, or other localization outputs for task-relevant objects as part of their reasoning trace. However, these approaches typically optimize final-answer correctness alone, allowing correct answers to be rewarded even when the model does not reason from confidently localized task-relevant objects. We introduce SpatialCORE (Spatially COnfident REasoning), a post-training framework that turns the model's own confidence in generated grounding into a learning signal for spatial reasoning. Its central idea is to reinforce grounding that is both accurate and confident, encouraging the model to reason from confidently localized task-relevant objects. SpatialCORE realizes this through a self-regulating spatial reward that weights each predicted bounding box's matching quality by its coordinate-token confidence. An answer gate further ties grounding optimization to final-answer correctness. SpatialCORE achieves state-of-the-art results among open-source and specialized spatial reasoning models across diverse benchmarks, and transfers effectively in zero-shot settings to unseen data distributions. The source code is available at https://github.com/rafiibnsultan/SpatialCORE.
Rafi Ibn Sultan, Xiangyu Zhou, Md. Sajid Alam Chowdhury +4
Department of Computer Science, Wayne State University · Department of Radiation Oncology, Henry Ford Health · Department of Electrical and Computer Engineering, The Ohio State University +1
Reliable spatial reasoning remains a core bottleneck for vision-language models (VLMs). Existing mainstream training paradigms for spatial reasoning largely rely on outcome alignment or process imitation, lacking explicit constraints on the reasoning process, and therefore struggle to ensure genuine visual dependence and stable reasoning trajectories. In this paper, we construct a high-quality CoT dataset covering diverse spatial phenomena and diagnose the model's reasoning process, revealing two typical types of process degradation during reinforcement learning optimization: Spurious Grounding, which bypasses visual evidence, and Tail Instability, where uncertainty abnormally rises in the later stage of reasoning. To address these issues, we propose ProSR, a process-shaping optimization framework for spatial reasoning. Through a Counterfactual Invariance Penalty and a Tail Drift Penalty, ProSR extends the optimization objective from single answer correctness to two process-level dimensions: visual dependence and trajectory stability. Experiments on multiple complex and out-of-distribution spatial reasoning benchmarks show that ProSR improves answer accuracy while generating reasoning trajectories that are more stable and more dependent on visual evidence.
Jiangyang Li, Cong Wan, Changjie Wu +8
1Xi’an Jiaotong University · 2Amap, Alibaba Group · 4Shenzhen University +1
Spatial VLMs have made substantial progress in geometric perception, yet complex spatial reasoning requiring multi-step inference over depth, distance, and scene relations remains challenging. Moreover, different spatial queries call for fundamentally different strategies: some are best addressed through purely linguistic, step-by-step deduction, while others require explicit 3D grounding before quantitative inference. We present Dual-Path Spatial Reasoning via Reinforcement Learning for Spatial VLMs (SR-REAL), a unified framework that equips a spatial VLM with two complementary reasoning paths: Language-Only Reasoning (LOR), which performs step-by-step linguistic deduction, and Detect-Then-Reason (DTR), which detects 3D geometric cues (e.g., centers or bounding boxes) via region tokens before explicit geometric inference. SR-REAL begins with a cold-start supervised fine-tuning stage that constructs LOR and DTR chain-of-thought supervision and exposes a region-to-3D interface, followed by RL that optimizes the policy model with accuracy and format rewards; for DTR, a discrete center-based detection reward further refines geometric alignment. Across diverse spatial benchmarks, SR-REAL significantly outperforms spatial VLM baselines: (i) a single RL-trained model supports both reasoning paths, with DTR excelling in region-aware tasks through precise 3D localization and LOR enhancing general spatial reasoning; (ii) jointly training both paths fosters mutual reinforcement; (iii) high-quality, blended cold-start data is crucial for stable RL optimization; and (iv) the model generalizes across datasets and domains without per-task tuning, demonstrating positive transfer between LOR and DTR.
Yatai Ji, An-Chieh Cheng, Yang Fu +13
The University of Hong Kong · NVIDIA · University of California, San Diego