Uruqi: Learning Spatial Cognition from Visual Experience
Organizations: Beijing University of Posts and Telecommunications · Tsinghua University
Abstract
Spatial intelligence requires maintaining a coherent understanding of the world as the embodied agent moves. Like humans, the agent must use its own motion to interpret changes across observations and update object locations and spatial relations accordingly. Despite spatial post-training having substantially broadened the spatial intelligence of vision-language models (VLMs), they still struggle with two atomic spatial capabilities: tracking self-motion and mapping the surrounding world during motion. To address this gap, we provide dense multi-turn supervision over interleaved atomic capabilities within each training episode, mimicking the visual experience of a continuously moving agent that reasons as it observes. To scale this up, we synthesize 11,738 motif-driven camera trajectories over a broad range of 3D scenes, supporting self-motion tracking, persistent object mapping, and rich spatial operations within each visual experience. By training models to reason over these atomic questions, our URUQI-8B improves accuracy from 15.84% to 47.73% on our Uruqi benchmark comprising 52k questions across 2.7k episodes. URUQI-SI-Mix-8B further reaches 50.41%, comparable to the 50.08% achieved by GPT-6 Astra. Trained solely on our synthesized data, URUQI-8B achieves an average relative accuracy improvement of 17.13% over its InternVL3-8B backbone across three external spatial benchmarks. These results highlight continuous visual experience as a scalable source of supervision for developing spatial cognition in VLMs.
Figures & tables
| Model | M1: self-motion | M2: object mapping | M3: state operations | Overall | ||||||||
|---|---|---|---|---|---|---|---|---|---|---|---|---|
| Mot. | Loc. | Yaw | ID | Pos. | Hist. | Inv. | Ref. | Upd. | Rel. | Evt. | ||
| Proprietary models | ||||||||||||
| Gemini 3.1 Pro | 13.64 | 17.96 | 43.58 | 29.41 | 11.84 | 25.15 | 48.65 | 25.40 | 34.82 | 27.37 | 25.56 | 27.37 |
| Qwen3.8-Max | 9.77 | 16.45 | 39.57 | 32.10 | 12.85 | 24.08 | 47.27 | 14.75 | 20.01 | 24.58 | 19.60 | 23.58 |
| GPT-5.5 | 11.86 | 21.16 | 44.53 | 31.22 | 11.36 | 20.18 | 44.20 | 22.43 | 19.04 | 21.30 | 20.90 | 24.50 |
| GPT-6 Astra | 25.31 | 41.30 | 89.70 | 52.53 | 29.72 | 54.91 | 61.38 | 47.13 | 63.25 | 45.89 | 37.68 | 50.08 |
| VSI-Bench | MMSI-Bench | MindCube-tiny | ||||||||
| Model | Num | MC | Avg | Pos. | MSR | Avg | Rot | Amg | Ard | Avg |
| Proprietary models (reference) | ||||||||||
| Gemini 3.1 Pro | 38.48 | 61.36 | 49.92 | – | – | 49.50 a | 90.50 | 71.33 | 83.20 | 77.81 |
| Qwen3.7-Plus | 61.56 | 68.51 | 65.04 | – | – | 45.00 a | 92.50 | 59.83 | 80.40 | 70.95 |
| GPT-5.5 | – | – | 60.40 a | – | – | 42.20 a | – | – | – | 65.50 a |
| GPT-6 Astra | 64.57 | 81.53 | 73.05 b | – | – | 57.90 c | – | – | – | 78.80 c |
| Model | QA history | Self-motion | Object mapping | |||||
|---|---|---|---|---|---|---|---|---|
| Translation error (m) | Heading error ( ∘ ) | Acc@0.5m (%) | Mean error (m) | |||||
| All | Initial | Visible | Absent | |||||
| InternVL3-8B | None | 0.397 | 93.57 | 1.00 | 0.00 | 2.42 | 0.00 | 3.004 |
| Model | 0.341 | 16.57 | 0.67 | 0.00 | 1.45 | 0.13 | 3.040 | |
| GT motion | 0.189 | 12.72 | 0.78 | 0.00 | 1.82 | 0.09 | 3.052 | |
| Qwen3.8-27B | None | 0.297 | 9.43 | 6.44 | 16.02 | 12.30 | 1.29 | 4.725 |
| Model | Method | Absent Acc@0.5m |
|---|---|---|
| Qwen3.8-27B | Direct prediction | 1.29 |
| Pose-based propagation | 16.88 | |
| URUQI Syn -8B | Direct prediction | 35.13 |
| Pose-based propagation | 71.61 |
Appendix figures & tables8 assets
Supplementary material from the paper’s appendix.
Appendix
| Operation family | Operation | Examples |
|---|---|---|
| Reference-frame reasoning | Express relations in a specified frame | Camera-relative direction; object-centered reference |
| Hypothetical motion | Update relations after imagined motion | Rotation; translation; turn–move sequences |
| Inverse spatial inference | Solve for an unknown from spatial constraints | Target identification from directional constraints |
| Measurement and comparison | Measure, compare, or rank geometry | Distance; size; height; area |
| Temporal and set operations | Query events and combine observations | Appearance order; reappearance; counts; set overlap |
| Route reasoning | Evaluate paths between locations | Route description; traversable-length comparison |
| Stage | Definition | Queries |
|---|---|---|
| Initial | Query at the object’s registration frame | 231 |
| Absent | Zero visible instance pixels | 2,240 |
| Reappeared | First observation with at least 256 instance pixels after an absence | 79 |
| Visible | Other observations with at least 256 instance pixels after registration | 1,651 |
| Weak visibility | Positive instance-pixel count below 256 | 6 |
| Total | 4,207 |
| Model | Visibility precision | Visibility recall | Pixel Hit | False-visible rate |
|---|---|---|---|---|
| InternVL3-8B | 54.13 | 88.87 | 7.52 | 66.12 |
| Qwen3.8-27B | 93.51 | 78.39 | 50.64 | 4.78 |
| URUQI Syn -8B | 68.31 | 99.49 | 71.68 | 40.54 |