Spatial intelligence requires maintaining a coherent understanding of the world as the embodied agent moves. Like humans, the agent must use its own motion to interpret changes across observations and update object locations and spatial relations accordingly. Despite spatial post-training having substantially broadened the spatial intelligence of vision-language models (VLMs), they still struggle with two atomic spatial capabilities: tracking self-motion and mapping the surrounding world during motion. To address this gap, we provide dense multi-turn supervision over interleaved atomic capabilities within each training episode, mimicking the visual experience of a continuously moving agent that reasons as it observes. To scale this up, we synthesize 11,738 motif-driven camera trajectories over a broad range of 3D scenes, supporting self-motion tracking, persistent object mapping, and rich spatial operations within each visual experience. By training models to reason over these atomic questions, our URUQISyn-8B improves accuracy from 15.84% to 47.73% on our Uruqi benchmark comprising 52k questions across 2.7k episodes. URUQI-SI-Mix-8B further reaches 50.41%, comparable to the 50.08% achieved by GPT-6 Astra. Trained solely on our synthesized data, URUQISyn-8B achieves an average relative accuracy improvement of 17.13% over its InternVL3-8B backbone across three external spatial benchmarks. These results highlight continuous visual experience as a scalable source of supervision for developing spatial cognition in VLMs.
Figures & tables
Figure 1: Visualization of two MMSI-Bench ( Yang et al., 2026c ) examples by querying M1 and M2 before asking the final question. The query order is shown below each plot. The coordinate frame is defined at C1 . The black and blue arrows mark the viewing directions at C1 and C2 , respectively, while the blue dashed line denotes the displacement from C1 to C2 predicted by M1. For M2, filled markers show object locations predicted from Image 1, while hollow markers show Image 2 predictions transformed into the C1 frame using the predicted self-motion. The overlap between filled and hollow markers reflects the cross-view consistency of object-location predictions.
Figure 2: Overview of Uruqi . Top: Visual experiences are generated by instantiating reusable motifs in 3D scenes, searching for feasible camera trajectories, and rendering RGB observations together with privileged depth and instance identities. Bottom: Each verified trajectory is compiled into supervision over three stages: self-motion tracking (M1), persistent object mapping (M2), and operations over spatial state (M3). The example illustrates spatial information evolving along a shared trajectory, with solid and dashed markers denoting visible and previously observed objects.
Model
M1: self-motion
M2: object mapping
M3: state operations
Overall
Mot.
Loc.
Yaw
ID
Pos.
Hist.
Inv.
Ref.
Upd.
Rel.
Evt.
↑
Proprietary models
Gemini 3.1 Pro
13.64
17.96
43.58
29.41
11.84
25.15
48.65
25.40
34.82
27.37
25.56
27.37
Qwen3.8-Max
9.77
16.45
39.57
32.10
12.85
24.08
47.27
14.75
20.01
24.58
19.60
23.58
GPT-5.5
11.86
21.16
44.53
31.22
11.36
20.18
44.20
22.43
19.04
21.30
20.90
24.50
GPT-6 Astra
25.31
41.30
89.70
52.53
29.72
54.91
61.38
47.13
63.25
45.89
37.68
50.08
Table 1: Evaluation on Uruqi benchmark. Each column reports task-specific accuracy (%). Subclasses are grouped by the three supervision stages. Overall is computed as the mean of the three stage scores, with subclasses weighted equally within each stage. Bold and underline denote the best and second-best scores, respectively.
VSI-Bench
MMSI-Bench
MindCube-tiny
Model
Num
MC
Avg
Pos.
MSR
Avg
Rot
Amg
Ard
Avg
Proprietary models (reference)
Gemini 3.1 Pro
38.48
61.36
49.92
–
–
49.50 a
90.50
71.33
83.20
77.81
Qwen3.7-Plus
61.56
68.51
65.04
–
–
45.00 a
92.50
59.83
80.40
70.95
GPT-5.5
–
–
60.40 a
–
–
42.20 a
–
–
–
65.50 a
GPT-6 Astra
64.57
81.53
73.05 b
–
–
57.90 c
–
–
–
78.80 c
Table 2: External spatial evaluation. Num and MC denote numerical and multiple-choice tasks in VSI-Bench; Pos and MSR denote positional and multi-step spatial reasoning in MMSI-Bench; Rot, Amg, and Ard denote rotation, among, and around in MindCube-tiny. Avg reports each benchmark’s overall score. Bold and underline indicate the best and second-best scores among open-source models, respectively. Superscripts indicate externally reported results: a SpatialAxiom ( Lou et al., 2026 ) ; b Wang ( Wang, 2026 ) ; c PhysBrain 1.5 ( Team et al., 2026 ) .
Model
QA history
Self-motion
Object mapping
Translation error (m) ↓
Heading error ( ∘ ) ↓
Acc@0.5m (%) ↑
Mean error (m) ↓
All
Initial
Visible
Absent
InternVL3-8B
None
0.397
93.57
1.00
0.00
2.42
0.00
3.004
Model
0.341
16.57
0.67
0.00
1.45
0.13
3.040
GT motion
0.189
12.72
0.78
0.00
1.82
0.09
3.052
Qwen3.8-27B
None
0.297
9.43
6.44
16.02
12.30
1.29
4.725
Table 3: Dense evaluation of self-motion estimation and egocentric object localization. Acc@0.5m measures the percentage of object predictions within 0.5 m of the ground-truth horizontal position. Bold and underline denote the best and second-best results per column, respectively.
Model
Method
Absent Acc@0.5m ↑
Qwen3.8-27B
Direct prediction
1.29
Pose-based propagation
16.88
URUQI Syn -8B
Direct prediction
35.13
Pose-based propagation
71.61
Table 4: Localization from visual estimates and camera poses. Acc@0.5m (%) is averaged over 2,240 absent-object queries. (a) Direct prediction uses each current VLM answer; pose-based propagation stores the first accepted object estimate and computes subsequent locations from camera poses. Both use No QA history, with GT poses supplied only to the geometry module. (b) Pose-based propagation uses Model-history object predictions; Predicted denotes poses integrated from the model’s motion estimates. Bold marks the best result within each model and panel. Complete metrics are in Table 7 .
Appendix figures & tables8 assets
Supplementary material from the paper’s appendix.
Appendix
Figure 3: Localization while the target remains visible. Under No QA history , Uruqi Syn -8B localizes the TV across changing viewpoints, with errors of 0.22, 0.03, and 0.03 m.
Figure 4: Localization after the target leaves view. Under No QA history , the picture is absent at the two later observations; Uruqi Syn -8B maintains low errors of 0.08 and 0.05 m in this selected trajectory.
Figure 5: Localization error during object absence. The room divider is absent at t=4 and t=16 . At t=16 , Uruqi Syn -8B has an error of 0.53 m under No QA history and 1.48 m under Model history. Top: No QA history ; bottom: Model history . Both conditions use the same object and observation steps, with identical coordinate limits across all six localization plots.
Figure 6: Localization after object reappearance. The painting is absent at t=27 and reappears at t=32 . For Uruqi Syn -8B, error decreases from 0.52 to 0.24 m under No QA history and from 0.80 to 0.70 m under Model history. Top: No QA history ; bottom: Model history . Both conditions use the same object and observation steps, with identical coordinate limits across all six localization plots.
Figure 7: Variation in the effect of answer history. For the absent cabinet, Uruqi Syn -8B has lower error under Model history at t=15 (0.41 versus 0.53 m), but higher error at t=26 (0.57 versus 0.07 m). Top: No QA history ; bottom: Model history . Both conditions use the same object and observation steps, with identical coordinate limits across all six localization plots.
Target identification from directional constraints
Measurement and comparison
Measure, compare, or rank geometry
Distance; size; height; area
Temporal and set operations
Query events and combine observations
Appearance order; reappearance; counts; set overlap
Route reasoning
Evaluate paths between locations
Route description; traversable-length comparison
Appendix
Table 5: Examples of spatial operations used in M3 supervision.
Stage
Definition
Queries
Initial
Query at the object’s registration frame
231
Absent
Zero visible instance pixels
2,240
Reappeared
First observation with at least 256 instance pixels after an absence
79
Visible
Other observations with at least 256 instance pixels after registration
1,651
Weak visibility
Positive instance-pixel count below 256
6
Total
4,207
Appendix
Table 6: Localization query stages in the dense evaluation. Pixel counts are normalized to 1024×1024 resolution; stages other than Initial apply after registration.
Model
Visibility precision ↑
Visibility recall ↑
Pixel Hit ↑
False-visible rate ↓
InternVL3-8B
54.13
88.87
7.52
66.12
Qwen3.8-27B
93.51
78.39
50.64
4.78
URUQI Syn -8B
68.31
99.49
71.68
40.54
Appendix
Table 8: Auxiliary visibility and object correspondence on 4,207 queries. All values are percentages; bold marks the best value in each column.
Vision-language models (VLMs) have shown strong performance on static visual understanding, yet they still struggle with dynamic spatial reasoning that requires imagining how scenes evolve under egocentric motion. Recent efforts address this limitation either by scaling spatial supervision with synthetic data or by coupling VLMs with world models at inference time. However, the former often lacks explicit modeling of motion-conditioned state transitions, while the latter incurs substantial computational overhead. In this work, we propose World2VLM, a training framework that distills spatial imagination from a generative world model into a vision-language model. Given an initial observation and a parameterized camera trajectory, we use a view-consistent world model to synthesize geometrically aligned future views and derive structured supervision for both forward (action-to-outcome) and inverse (outcome-to-action) spatial reasoning. We post-train the VLM with a two-stage recipe on a compact dataset generated by this pipeline and evaluate it on multiple spatial reasoning benchmarks. World2VLM delivers consistent improvements over the base model across diverse benchmarks, including SAT-Real, SAT-Synthesized, VSI-Bench, and MindCube. It also outperforms the test-time world-model-coupled methods while eliminating the need for expensive inference-time generation. Our results suggest that world models can serve not only as inference-time tools, but also as effective training-time teachers, enabling VLMs to internalize spatial imagination in a scalable and efficient manner.
Wanyue Zhang, Wenxiang Wu, Wang Xu +6
Institute of Automation, Chinese Academy of Sciences · School of Artificial Intelligence, University of Chinese Academy of Sciences · Harbin Institute of Technology +2
Vision-Language Models (VLMs) perform well on commonsense reasoning tasks but struggle with visual spatial reasoning. Most existing solutions introduce extra 3D prior inputs or external spatial encoders, which increase complexity and degrade the underlying VLMs' general-purpose capabilities after spatial fine-tuning. To this end, we propose a parameter-efficient \textit{\textbf{Spatio}-vision \textbf{L}anguage \textbf{M}odels (SpatioLM)}, that enhances spatial intelligence without extra 3D prior inputs or third-party spatial encoders. Concretely, we design a plug-and-play and non-invasive spatio-vision module that elicits the spatial knowledge inherent in VLMs. Furthermore, we innovatively leverage pseudo depth and camera information as supervision to guide the model in learning physically coherent representations. Extensive experiments show that SpatioLM achieves significant improvements in diverse tasks, including spatial perception and understanding while effectively limiting the degradation of general capabilities. Notably, the model achieves an impressive score of 71.6 on the VSI-Bench (the first model to surpass 70). In addition, it attains competitive performance when transferred to embodied manipulation tasks. Code is available at \faGithub~spatio-lm.
Jing Wu, Jianhua Wu, Jiayi Guan +5
Xiaomi EV, Beijing, China · College of Automotive and Energy Engineering, Tongji University, Shanghai, China · Independent Researcher
We push the frontier of large-scale spatial intelligence in Vision-Language Models (VLMs) and introduce the first benchmark that probes geographical layout understanding from real-world videos, spanning up to 1km distances. Inspired by the cognitive science literature, we evaluate models against the hierarchical stages of human spatial awareness: anchoring via landmarks, connecting them through routes, and integrating these into global mental maps. Extensive experiments reveal a fundamental divergence in how current AI models process spatial information. Instead of utilising true path integration or forming geometric survey knowledge, we find that VLMs rely almost entirely on 2D visual recognition and text-matching to bypass complex spatial reasoning. The benchmark is publicly available at https://perception-test-challenge.github.io/kilometervision.html.
Aravindh Mahendran, Michael King, Matthew Koichi Grimes +12
Google DeepMind, Berlin, Germany · Google DeepMind, London, UK · Princeton University, Princeton, USA +3