Authors: Rafi Ibn Sultan, Md. Sajid Alam Chowdhury, Saleh Zare Zade, Chengyin Li, Prashant Khanduri, Marco Brocanelli, Dongxiao Zhu
Organizations: Department of Computer Science, Wayne State University · Department of Radiation Oncology, Henry Ford Health · Department of Electrical and Computer Engineering, The Ohio State University · Institute for AI and Data Science, Wayne State University
Large Vision-Language Models (LVLMs) commonly perform spatial reasoning through chain-of-thought (CoT), encoding intermediate reasoning as autoregressive sequences of discrete language tokens. Such hard thinking requires committing to a single token at each step, even when the correct spatial interpretation remains uncertain. This early commitment constitutes premature discretization: an incorrect token selection can propagate errors through subsequent reasoning. We propose Soft Spatial Reasoning, a post-training framework that introduces soft thinking for spatial tasks in LVLMs. At each intermediate reasoning step, the LVLM forms a continuous soft state by mixing token embeddings rather than selecting a single token, allowing multiple candidate continuations to influence the next step. The appropriate degree of softness, however, can vary across reasoning steps: retaining multiple candidates may preserve a useful spatial interpretation, but if those candidates imply conflicting spatial relations, mixing them may interfere with subsequent reasoning. At the core of Soft Spatial Reasoning is AdaptSoft, a controller that uses the current hidden state and predictive uncertainty to adapt the degree of softness at each reasoning step. To train AdaptSoft, we introduce a gradient-alignment learning objective that provides a step-specific learning signal for softness control without intermediate reasoning supervision. Across diverse spatial benchmarks, Soft Spatial Reasoning outperforms hard and fixed-soft CoT baselines using the same backbone, as well as a range of existing LVLMs. The source code is available at https://github.com/rafiibnsultan/Soft_Spatial_Reasoning
Figures & tables
Figure 1: Hard versus adaptive soft thinking for spatial reasoning. (a) Hard thinking commits to one token at each CoT step, allowing early errors to propagate. (b) Our Soft Spatial Reasoning forms continuous intermediate states by mixing token embeddings, carrying information from multiple candidate continuations through the CoT. AdaptSoft adjusts the mixture’s softness at each step based on the current reasoning state and predictive uncertainty.
Figure 2: Overview of Soft Spatial Reasoning . (a) The LVLM policy is post-trained to carry multiple candidate continuations through soft states during reasoning, then generate the final answer as discrete tokens. Rollout rewards yield group-relative advantages for the GRPO update. (b) At each reasoning step, AdaptSoft sets the mixture temperature from the current hidden state and candidate-distribution entropy. The resulting mixture weights combine token embeddings into the next soft state.
Figure 3: Gradient-Alignment Learning. At each soft reasoning step t , the temperature τi,t controls the state si,t , whose contribution gi,t to the LVLM output-layer gradient is compared with a reference gradient Gref computed from a disjoint subset of rollouts. The resulting alignment scores αi,t are aggregated to form the AdaptSoft loss.
Method
Average
Dynamic Reasoning
Spatial Interaction
Complex Logic
Perspective Taking
Mani- pulation
Motion Anal.
Traffic Anal.
Loca- lization
Geospa. Strategy
Pattern Rec.
Geometric Reasoning
Ego Centric
Allo Centric
Hypo- thetical
Reference Baselines
Random Choice
24.98
24.86
26.30
25.88
23.43
27.27
21.44
24.77
22.55
24.84
25.78
Human Evaluation
92.63
94.62
96.07
91.38
95.11
92.15
89.02
85.90
98.53
94.30
90.26
Proprietary Models
GPT-4.1-mini Open (2025)
48.87
64.32
56.53
59.06
60.19
56.36
29.28
30.19
72.55
39.57
39.28
Table 1: OmniSpatial Jia et al. (2025) results across 10 spatial reasoning task categories. Soft Spatial Reasoning is compared against proprietary models, general open-source LVLMs, soft thinking models, and specialized spatial reasoning models. Proprietary models are included as reference points; the best open-source or specialized result is shown in bold . Average accuracy is weighted by category sample size.
Figure 4: Soft Spatial Reasoning on a randomly selected example from OmniSpatial. AdaptSoft varies softness across the CoT. For readability, greedy decoding is used only to display the soft states as text. Shading shows each span’s mean temperature ( blue : lower; orange : higher). The plot below shows temperature at every reasoning step and labels several words with their temperatures.
Method
Average
Question Categories
3D Geom.
Dep. & Occu.
Orientation
Relat. Posit.
Size & Scale
Spati. Navig.
Reference Baselines
Random Choice
25.00
25.00
25.00
25.00
25.00
25.00
25.00
Human Baseline
87.57
93.70
74.13
91.58
91.51
88.89
87.76
Proprietary Models
GPT-4o-mini Hurst et al. (2024)
46.50
47.06
39.00
47.03
47.17
49.60
49.79
Table 2: SpatiaLab Wasi et al. (2026) results across 6 spatial reasoning task categories in the zero-shot setting. Proprietary models are included as reference points; the best open-source or specialized result is shown in bold . Average accuracy is weighted by category sample size.
Method
Overall
Rotation
Among
Around
Reference Baseline
Random Choice
32.35
36.36
32.29
30.66
Proprietary Models
Gemini-2.5-Pro Team et al. (2023)
47.05
85.50
25.95
38.40
Claude-4-Sonnet Anthropic (2024)
44.75
48.42
44.21
47.62
Open-weight Models
Table 3: Zero-shot results on MindCube Wang et al. (2026a) across three spatial mental-modeling settings. Proprietary models are included as reference points, and Overall is weighted by setting size.
Variant
Avg.
Ego centric
Allo centric
Hypo thetical
Soft Spatial Reasoning (Full)
42.75
81.37
32.53
41.46
w/o hi,t in AdaptSoft
39.22
77.45
29.52
36.14
w/o predictive uncertainty
40.64
78.43
30.85
38.55
w/o gradient alignment
38.68
76.47
29.26
34.94
w/o gradient centering
41.18
79.41
31.12
39.76
Table 4: OmniSpatial Perspective Taking ablations Jia et al. (2025) . All variants share the backbone and policy GRPO setup, with post-training and evaluation on the corresponding training and test splits. Averages are sample-weighted.
Appendix figures & tables5 assets
Supplementary material from the paper’s appendix.
Appendix
Figure 5: AdaptSoft Temperatures During Training. The line shows the mean temperature τi,t across soft steps in each rollout batch; shading spans their minimum and maximum. The dashed line marks τ0=0.5 .
Figure 6: AdaptSoft Temperature and Predictive Uncertainty. Temperatures and predictive uncertainty are measured across soft reasoning steps during inference on the test set. Color shows the number of steps on a logarithmic scale. The solid line shows median temperature in each entropy bin; dashed lines mark the 10th and 90th percentiles.
Figure 7: AdaptSoft Varies Softness Within CoTs. The histogram shows the 90th–10th percentile temperature spread within each test-set reasoning trace generated during inference. The solid line marks the median within-trace spread; the dashed line marks the corresponding spread across trace-median temperatures.
Hyperparameter
Value
Notes
Model configuration
Backbone
Qwen3-VL-8B-Thinking
–
Language-model adaptation
Full fine-tuning
No LoRA
Numerical precision
BF16
FSDP2 mixed precision
Parameter sharding
FSDP2
Parameter and optimizer offload
Data preparation
Appendix
Table 5: Post-training and evaluation configuration for Soft Spatial Reasoning.
Figure 8: System prompt used to standardize the reasoning format during Soft Spatial Reasoning training.
Department of Computer Science, Wayne State University · Department of Radiation Oncology, Henry Ford Health · Department of Electrical and Computer Engineering, The Ohio State University +1