Despite progress in vision-language models, 3D spatial reasoning from 2D images remains challenging. Text-based methods describe intermediate geometry with discrete tokens, limiting fidelity for continuous spatial relations. Continuous latents offer richer representations, but a single latent type does not explicitly separate the cues needed across spatial tasks. Decomposed spatial latents address this by representing position, direction, and global geometry separately under geometric supervision. Yet the geometry representation can still collapse toward one dominant direction, and unrestricted attention can leave the latents underused during answer learning. We introduce GeoLatent, combining Common--Residual Geometry Alignment (CR-GEO) with routed optimization to structure the geometry states while promoting latent-mediated answer learning. CR-GEO separates shared from residual teacher geometry; routed optimization jointly trains geometry and language, temporarily directs visual answer learning through the latents, and restores full attention with geometry supervision. In controlled comparisons, CR-GEO raises geometry effective rank from 1.00 to 3.87, while blocking latent readout at the bottleneck lowers direction accuracy from 89.1% to 25.8% on 128 fixed questions. After recovery, the differentiated geometry representation and latent-mediated visual route remain available alongside direct image access. GeoLatent achieves 73.0% on SPAR-Bench and 72.1% on SPBench, outperforming previously reported methods on both.
Figures & tables
Figure 1: GeoLatent addresses two limitations of decomposed latent reasoning: CR-GEO reduces redundancy among geometry states, while routed optimization encourages answer learning to use the decomposed spatial latents.
Figure 2: GeoLatent overview. (a) A VLM interleaves language with decomposed POS, DIR, and GEO latents. (b) Routed optimization proceeds through joint, bottleneck, and recovery stages, temporarily restricting direct visual access during answer learning. (c) CR-GEO separates shared from residual teacher geometry for GEO supervision.
Figure 3: GEO-token organization. The public checkpoint and models trained with the original GEO alignment loss approach rank one with near-zero assignment information, whereas CR-GEO yields differentiated GEO vectors. The joint pair comes from the matched-initialization loss-only comparison; the recovered pair comes from separately trained recovered controls in Table 2 .
Model
Params.
SPAR-Bench
SPBench
ViewSpatial
Mean
Avg
Dep
Dis
Prox
Rel
View
Avg
Rel
Abs
Avg
Proprietary models reported by Li et al. (2026c)
GPT-4o † ( Hurst et al., 2024 )
–
40.1
34.3
43.4
57.1
45.9
31.0
53.4
49.4
56.0
37.5
43.7
Gemini-2.5-Flash † ( Comanici et al., 2025 )
–
48.7
37.4
45.2
80.1
63.2
41.0
51.5
44.7
56.0
44.0
48.1
General VLMs reported by Li et al. (2026c)
MiniCPM-V-4.5 † ( Yu et al., 2026 )
8B
37.7
32.4
32.6
62.1
50.6
29.8
40.9
47.4
36.7
39.0
39.2
Table 1: Spatial-reasoning performance (%). † : results reported by Li et al. (2026c) ; SPAR numerical entries use threshold-averaged accuracy; others use accuracy. Mean averages the three benchmarks. Best and second-best entries are bold and underlined, respectively.
Figure 4: Latent use and sample-specific visual dependence. (a) Blocking downstream readout from all decomposed spatial latents leaves public GeoAnchor unchanged but reduces the matched K=4/8 bottleneck models to near-chance accuracy. (b) ViewSpatial under real (filled), blank (open), and mismatched images (cross), including the matched no-GEO control; VEF is the fraction of above-chance accuracy removed by blanking.
Appendix figures & tables13 assets
Supplementary material from the paper’s appendix.
Appendix
Source
Task group
Answer format
Count
ScanNet / SPAR
depth_prediction_oc
numeric
4,000
depth_prediction_oo
numeric
4,000
distance_prediction_oc
numeric
4,000
distance_prediction_oo
numeric
4,000
distance_infer_center_oo
select
8,000
obj_spatial_relation_oo
select
8,000
Appendix
Table 4: Spatial reasoning training mixture. “Select” denotes multiple-choice answers and “numeric” denotes fill-in numerical answers. Structured3D spatial-imagination is sentence-format in the released SPAR records.
Benchmark
Reported group
Underlying task(s)
Count
SPAR-Bench
Dep.
depth_prediction_oc, depth_prediction_oo
732
Dis.
distance_prediction_oc, distance_prediction_oo
756
Prox.
distance_infer_center_oo
340
Rel.
obj_spatial_relation
364
View
spatial_imagination_oc, spatial_imagination_oo
674
SPBench
Rel.
object_rel_direction, object_rel_distance
397
Appendix
Table 5: Benchmark task groups and sample counts. SPAR-Bench follows the single-image subset used in GeoAnchor. ViewSpatial additionally tests transfer from camera-frame supervision to camera- and person-centered reference frames.
Benchmark
Record ID
Scene
Scene/frame
Image SHA
Question
Exact example
SPAR-Bench
0
–
–
0
63
0
SPBench
0
23
5
0
777
0
ViewSpatial
0
14
2
0
0
0
Appendix
Table 6: Overlap of each evaluation set with the 105,928-record training mixture. Entries count affected benchmark records. “Exact example” requires matching image or scene/frame, normalized question, and answer. A dash means the benchmark does not expose that metadata.
Task
N
Letter max
Semantic max
Direction
Camera–relative direction
1,773
26.4
16.0
right
Camera–object orientation
996
26.0
26.8
back
Person–object orientation
996
26.1
57.2
front
Person–relative direction
842
34.8
27.4
right
Appendix
Table 8: ViewSpatial answer priors. “Letter max” is the frequency of the most common correct option position; “semantic max” is the most frequent direction irrespective of option position.
Phase
Steps
LR
Attention policy
λNTP
λlocal
λGEO
Joint
3,311
2×10−5
full
1.0
1.0
0.2
Bottleneck
1,000
1×10−5
decomposed-latent bottleneck
1.0
0.1
0.1
Recovery
3,311
2×10−5
full
1.0
0.1
0.2
Appendix
Table 9: GeoLatent training phases. All phases use Qwen3-VL-2B-Instruct, frozen vision encoder, trainable merger/LLM/latent modules, global batch 32, bf16 computation, and fp32 optimizer states.
Table 10: Full loss and curriculum matrix. The post-joint recovery variants use local/GEO supervision weights 0/0 (NTP-only), 0.1/0 (local-only), 0.01/0.05 , and 0.1/0.2 (standard GeoLatent recovery). “Reference weights” denotes the stage-specific weights of the original- Lcov reference configuration, while “GeoLatent weights” denotes GeoLatent’s stage-specific local/GEO supervision weights.
K
Joint
Bneck
SPAR
SPB
VS
Rec. mean
reff
Samples/s
2
61.2
56.5
72.8
71.9
43.9
62.9
1.70
3.37
4
61.3
53.7
74.1
71.9
45.0
63.6
3.23
2.86
8
60.4
55.9
72.8
72.0
42.5
62.4
3.81
2.46
16
60.7
52.9
73.2
72.2
42.1
62.5
4.05
1.77
Appendix
Table 11: GEO-vector count ablation under matched full-curriculum runs. Stage columns report the three-benchmark mean; recovery columns give individual benchmark scores. Throughput is measured during recovery.
GEO objective
SPAR
SPBench
ViewSpatial
Mean
reff
Iassign
original Lcov
67.90
66.14
40.53
58.19
1.00
≈0
CR-GEO LCR-GEO
68.63
67.83
38.83
58.43
3.84
0.178
Appendix
Table 12: Matched cross-backbone reproduction with Qwen2.5-VL-3B. Benchmark entries are accuracy/score (%); structure is measured on the same frozen 768-frame protocol used for the main backbone. Mean is the unweighted average of the three benchmark scores.
Figure 5: Qualitative structure of the learned spatial latents. The t-SNE projection illustrates reduced concentration of GEO vectors under CR-GEO, and the local-token maps show object-specific visual attention during latent generation. Quantitative structure and use are reported in Table 14 and Figure 4 .
Objective
Stage
TF POS ↓
TF DIR ↑
Composed DIR ↑
Live POS ↓
Live DIR ↑
Live/replay DIR agreement ↑
CR-GEO
joint
0.347
0.923
0.871
0.307
0.942
1.000
CR-GEO
bottleneck
0.365
0.911
0.864
0.345
0.919
0.971
CR-GEO
recovery
0.338
0.899
0.874
0.300
0.920
1.000
original Lcov
joint
0.354
0.890
0.866
0.310
0.909
1.000
original Lcov
bottleneck
0.376
0.876
0.862
0.345
0.884
0.957
original Lcov
recovery
0.335
0.907
0.868
0.305
0.916
1.000
Appendix
Table 13: POS/DIR decoding across the matched 2×3 objective–stage grid. TF denotes teacher-forced evaluation. POS is mean per-point L2 error in meters; DIR and composed DIR are cosine similarities to ground truth. Live/replay DIR agreement is the cosine agreement between the DIR representation captured during live generation and the DIR representation obtained by teacher-forced replay of the same generated sequence. Live evaluation starts from the fixed 192-record subset; reported live metrics use the successfully aligned cases.
Model
off-diag cos
reff
Hcond/logK
Iassign
NLSE
original Lcov , matched-init joint
1.000
1.00
–
≈0
–
CR-GEO, matched-init joint
0.167
3.87
–
0.228
–
GeoAnchor public
0.991
1.05
0.999
≈0
7.54
original Lcov , joint (reference weights)
0.460
2.73
0.667
≈0
4.00
original Lcov , recovery (reference weights)
1.000
1.00
1.000
≈0
7.99
original Lcov , joint (GeoLatent weights)
1.000
1.00
–
≈0
–
Appendix
Table 14: GEO-token representation diagnostics on frozen samples ( n=768 ). reff is the effective rank of the 8×8 cosine Gram matrix. Iassign is normalized information induced by teacher-to-vector assignments, not a general neural mutual-information estimator. NLSE is the effective number of covering vectors. Dashes denote diagnostics unavailable from historical captures, not failed measurements.
K=4
K=8
Condition
Acc.
Drop
Acc.
Drop
No latent-readout cut
85.16
–
89.06
–
Block POS
61.72
23.44
53.12
35.94
Block DIR
81.25
3.91
85.94
3.13
Block POS+DIR
26.56
58.59
26.56
62.50
Block GEO
85.16
0.00
89.06
0.00
Appendix
Table 15: Decomposed-latent localization under bottleneck inference. Entries are accuracy (drop from the no-cut condition), in percent. “Block prior-latent access” prevents all subsequent typed and non-typed queries from attending to previously generated typed-latent keys. For both K=4 and K=8 , neither the POS+DIR drop nor the latent-readout drop was matched or exceeded by any of 20,000 two-sided scene-clustered permutations (plus-one p=1/20001 for each).
Checkpoint
No latent-readout cut
Block latent readout
Drop
Joint (full attention)
69.53
68.75
0.78
Bottleneck (routed)
89.06
25.78
63.28
Recovery (full attention)
82.03
81.25
0.78
Appendix
Table 16: Stage-wise route intervention in the matched K=8 trajectory on the same 128 directions. “Block latent readout” hides all decomposed-latent keys from subsequent non-latent queries while preserving latent-to-latent relay.
Although multimodal large language models (MLLMs) have achieved remarkable progress, understanding 3D spatial relationships from 2D images remains a critical challenge. Existing methods primarily rely on symbolic text tokens, which inherently lack the fidelity to represent continuous geometric information. While recent methods use latent representations to enhance reasoning, relying on a single latent type cannot adapt to the diversity of spatial tasks, leading to misalignment in complex geometric scenarios. To address these limitations, we propose GeoAnchor, an interleaved text-latent reasoning framework. GeoAnchor decomposes 3D spatial information into three complementary components: position latents for object grounding, direction latents for relational orientation, and geometry latents for scene structure. These components are recombined in a structured space to construct local evidence while capturing global context, enabling dynamic and interpretable reasoning. Furthermore, we introduce a collaborative training strategy that guides the model from local spatial perception to comprehensive 3D understanding. Extensive experiments on diverse and complex 3D reasoning tasks demonstrate that GeoAnchor outperforms the state of the art, validating its effectiveness and generalization capabilities.
Hao Li, Han Fang, Zixin Pan +8
Shanghai Jiao Tong University · Xingchen AGI Lab, China Telecom Artificial Intelligence Technology (Beijing) · The Hong Kong University of Science and Technology (Guangzhou) +1
Vision-Language Models (VLMs) often struggle with robust 3D spatial reasoning. Prevailing methods that rely on fine-tuning with 3D visual question-answering (VQA) datasets may overfit dataset-specific biases, while integrating specialized 3D visual encoders is often inflexible and cumbersome. In this paper, we argue that genuine spatial understanding should emerge from learning fundamental geometric priors, not only from high-level VQA supervision. We propose GASP (Geometric-Aware Spatial Priors), a framework that injects these priors directly into the LLM's transformer layers. GASP employs a small correspondence head, applied as a deep supervision signal across all layers, and is trained with a dual objective leveraging ground-truth geometry from large-scale video scenes: a contrastive loss on ground-truth point correspondences enforces 2D view-invariance, while a depth consistency supervision resolves 3D geometric ambiguities. Our analysis first provides a diagnostic showing that standard VLMs' internal correspondence matching accuracy is very low (often below 5%). We then demonstrate that our training substantially improves this behavior, boosting peak layer-wise correspondence to over 70% and maintaining over 85% temporal robustness while baselines remain below 5%. These internal improvements translate to significant gains on downstream spatial benchmarks including +18.2% on All-Angles Bench and +29.0% on VSI-Bench, all without training on any 3D VQA data. Our findings indicate that learning from fundamental geometric priors is a promising and generalizable pathway towards VLMs with more reliable 3D spatial reasoning.
Vision-language models (VLMs) can benefit from geometric priors for multi-view spatial reasoning, yet answer-only training does not directly supervise the intermediate geometric estimates and their use in deriving quantitative spatial answers. We hypothesize that spatial chain-of-thought (CoT) supervision becomes more effective when the VLM first jointly learns complementary local geometry and global scene context through multi-view reconstruction. We introduce SpatialSpeak, a two-stage framework that connects QA-native reconstruction pretraining with spatial CoT learning. In Stage I, QA-Native Reconstruction Pretraining (QA-RP) combines marked-point 3D queries for fine-grained local geometry with object-center queries for global scene context across views. Both tasks are formulated as text-based question answering, allowing geometric estimation and subsequent reasoning to share the same autoregressive output interface. In Stage II, spatial CoT with Visual Compensation (CoT-VC) trains the model to express question-relevant geometric estimates and use them to derive answers, with reliability assessment and visual compensation supporting answer refinement when needed. On ReVSI, QA-RP increases the gain from CoT-VC from 2.6 to 6.9 points, and ablations show that both local and global reconstruction supervision are beneficial. SpatialSpeak achieves state-of-the-art results on ReVSI, VSI-Bench, and SPAR-Bench, with a ReVSI score of 62.8 that exceeds the strongest compared baseline by 8.7 points.
Yang Cao, Jiaxin Zhang, Dave Zhenyu Chen +4
Hong Kong University of Science and Technology · Harbin Institute of Technology · Huawei Noah’s Ark Lab