Existing spatial reasoning benchmarks mainly test spatial perception: reading off relations already visible in the input. Yet real-world spatial intelligence demands predictive spatial reasoning: constructing a scene from observations, anticipating how an intervention changes it, and reasoning about the unseen outcome. We introduce SpaceCast-Bench, the first benchmark to directly and diagnostically evaluate this capability. Built around an observe-transform-infer framework, its 3,862 questions from 182 real-world scenes span 16 task types at three levels: static perception, local prediction, and global prediction, progressively requiring scene understanding, spatial state updating, and relational inference over unobserved outcomes. Evaluating 21 models exposes a stark gap: the strongest model reaches only 58.0% against 87.2% human performance, while spatially specialized models remain near random chance. Controlled analyses further reveal that bridge views are critical for integrating distributed observations, and that explicit 3D evidence benefits models more reliably than generated outcome images or videos. Fine-tuning on our programmatically generated data lifts Qwen3-VL-4B from 34.0% to 65.7% with macro-average gains across six out-of-domain benchmarks.
Figures & tables
Figure 1: We introduce SpaceCast-Bench, a benchmark for evaluating predictive spatial reasoning. Left: representative tasks at three capability levels. Right: task-wise performance of six representative models, with human and random baselines.
Figure 2: Overview of the SpaceCast-Bench construction pipeline.
Figure 3: Representative tasks of SpaceCast-Bench. Correct answers are highlighted in green.
Figure 4: Task distribution of SpaceCast-Bench.
Cam. Dir.
Obj. Dir.
World Dir.
Dist.
Occ.
Att.
Avg.
Move Cam. Dir.
Move Obj. Dir.
Rotate Obj. Dir.
Move World Dir.
Move Dist.
Move Occ.
Remove Occ.
Avg.
Rotate Cam. Dir.
Rotate Obj. Dir.
Rotate World Dir.
Avg.
Model
Rank
Overall
L1: Static
L2: Local
L3: Global
Baseline
Human Level
–
87.2
92.0
92.4
90.1
89.7
100.0
97.9
93.7
81.6
83.2
83.6
76.1
82.5
84.2
97.2
84.1
81.0
82.4
80.8
81.4
Random Level
–
25.9
25.8
25.9
24.0
25.7
29.9
11.3
23.8
23.5
21.1
24.2
23.1
21.8
37.4
38.0
27.0
24.3
31.6
26.3
27.4
Proprietary Models
Claude Sonnet 4.6
9
36.0
29.5
17.1
35.0
40.1
71.5
66.0
43.2
38.2
14.5
30.1
30.3
30.5
43.9
40.7
32.6
39.9
15.6
33.8
29.8
Table 1: Accuracy (%) on all 16 SpaceCast-Bench task types. “Cam.”, “Obj.”, and “World” denote camera-centric, object-centric, and world-centric reference frames; “Dir.”, “Dist.”, “Occ.”, and “Att.” denote direction, distance, visibility, and attachment. “Move”, “Rotate”, and “Remove” denote local intervention types. Bold indicates the best model result in each column; baselines are excluded.
Table 6Figure 7
Figure 7: Effect of external spatial evidence. Left: Mean accuracy change (pp) relative to baseline across three model categories (proprietary, open-source, and spatial) when augmented with BEV images, textual 3D metadata, or generated outcome evidence from Qwen-Image-Edit, GPT-image-2, and Seedance2-mini. Right: A representative failure where the reference object is severely deformed in the generated outcome image.
Appendix figures & tables11 assets
Supplementary material from the paper’s appendix.
Appendix
Figure 8: Distribution of retained bridge views across the 3,328 cross-view questions.
Task
Question Template
Camera Direction
• From the first main view’s camera perspective, { object 1 } is in which direction relative to { object 2 } ? • Using the first main view’s camera coordinate frame, where is { object 1 } positioned relative to { object 2 } ? • From the first main view’s camera perspective, what is the spatial relationship of { object 1 } to { object 2 } ?
Object Direction
• Imagine you are { object 1 } and facing toward { object 2 } . From your perspective, in which direction is { object 3 } ? • If you were { object 1 } , looking toward { object 2 } , where would { object 3 } be?
World Direction
• The first main view’s camera is facing { camera direction } . On the room’s floor plan, in which cardinal direction is { object 1 } from { object 2 } ? • In the first main view, the camera faces { camera direction } . Viewed from above on the room’s layout, { object 1 } is in which cardinal direction relative to { object 2 } ?
Distance
• What is the approximate shortest distance between { object 1 } and { object 2 } , measured from their closest points? • Measured from the closest points of each object, what is the approximate shortest distance between { object 1 } and { object 2 } ? • Estimate the approximate shortest distance between { object 1 } and { object 2 } , measured from their closest points.
Occlusion
• What is the occlusion status of { object 1 } in the current view? • From the current viewpoint, which best describes { object 1 } : not occluded, occluded, or not visible? • In the current image, is { object 1 } unoccluded, occluded by another object, or not visible?
Attachment
• Suppose { object 1 } were moved to a different location. Which of the following objects would also be displaced from their current positions? • If { object 1 } were relocated elsewhere in the room, which of the following objects would also change position? One or more options may be correct. Select all that apply. • Imagine { object 1 } is moved to a new spot. Which of the following objects would also be displaced as a result? One or more options may be correct. Select all that apply.
Appendix
Table 6: L1 (Static Perception) templates for six task types querying original-scene relations or attachment structure. Red placeholders mark variable content.
Task
Question Template
Move Camera Direction
• From the first main view’s camera perspective, imagine moving { object 1 } { direction } by { distance } . After this change, what is the relative position of { object 2 } to { object 3 } ? • From the first main view’s camera perspective, if we move { object 1 } { direction } by { distance } , where is { object 2 } relative to { object 3 } ? • From the first main view’s camera perspective, suppose { object 1 } is shifted { direction } by { distance } . After this change, what is the spatial relation of { object 2 } with respect to { object 3 } ?
Move Object Direction
• In the initial scene, imagine { object 1 } faces the camera in the first main view, and you are { object 2 } , also facing that same camera. Freeze both objects’ initial horizontal forward/right axes before anything moves. If { object 1 } is moved { direction } by { distance } in its own frozen facing frame, from your separately frozen facing frame, in which horizontal direction is { object 3 } from you afterward? • Suppose { object 1 } and { object 2 } each initially face the camera in the first main view. Keep each object’s initial horizontal facing axes fixed throughout the motion. After moving { object 1 } { direction } by { distance } in { object 1 } ’s frozen frame, where is { object 3 } relative to { object 2 } in { object 2 } ’s separately frozen frame?
Rotate Object Direction
• Imagine you are { object 1 } and facing toward { object 2 } . If { object 3 } were moved along a { angle } -degree { rotation direction } (viewed from above) orbit around the center of { object 2 } in the horizontal plane, without changing its own facing direction, from your perspective, in which direction would { object 4 } be?
Move World Direction
• If { object 1 } is moved { distance } to the { direction } , the camera in the first main view faces { camera direction } . On the floor plan, in which cardinal direction would { object 2 } be from { object 3 } ? • After moving { object 1 } { distance } to the { direction } , on the room’s layout (the first main view’s camera faces { camera direction } ), in which cardinal direction is { object 2 } from { object 3 } ?
Move Distance
• From the first main view’s camera perspective, imagine moving { object 1 } { direction } by { distance } . After this change, what is the approximate shortest distance between { object 2 } and { object 3 } , measured from their closest points? • From the first main view’s camera perspective, if { object 1 } is moved { direction } by { distance } , what is the approximate shortest distance between { object 2 } and { object 3 } , measured from their closest points?
Move Occlusion
• After moving { object 1 } { direction } by { distance } , which best describes the pairwise occlusion relationship between { object 2 } and { object 3 } from the last main view’s viewpoint? • From the last main view’s viewpoint, after { object 1 } is moved { direction } by { distance } , is { object 2 } occluded by { object 3 } , is { object 3 } occluded by { object 2 } , or does neither object occlude the other?
Appendix
Table 7: L2 (Local Prediction) templates for seven task types involving object translation, orbit, or removal. Red placeholders mark variable content.
Task
Question Template
Rotate Camera Direction
• Suppose this room had originally been designed with its orientation rotated { angle } degrees clockwise about the camera position in the first photo below (viewed from above), with every object keeping its position relative to the others. Observed from that same camera position and viewing direction (unchanged), in which direction is { object 1 } relative to { object 2 } ? • If the room layout had been rotated { angle } degrees clockwise about the camera position in the first photo below (top-down view) from the start, with every object keeping its position relative to the others and that camera’s position and orientation unchanged, from that camera’s perspective, where would { object 1 } be relative to { object 2 } ? • Imagine the room was originally built rotated { angle } degrees clockwise about the camera position in the first photo below (as seen from above). With every object keeping its position relative to the others and that same camera at its original pose, from that camera’s perspective, what is the direction of { object 1 } from { object 2 } ?
Rotate Object Direction
• Suppose this room had originally been oriented { angle } degrees clockwise about the camera position in the first anchor image (viewed from above), with every object keeping its position relative to the others. If you were { object 1 } at its rotated position and kept facing the same horizontal direction that originally pointed from { object 1 } toward { object 2 } , in which direction would { object 3 } be?
Rotate World Direction
• Imagine all furniture is rotated { angle } degrees clockwise about the camera position in the first photo below (viewed from above), with every object keeping its position relative to the others. That same camera, facing { camera direction } , remains in place. On the floor plan, in which cardinal direction is { object 1 } from { object 2 } ?
Appendix
Table 8: L3 (Global Prediction) templates for three direction tasks after a counterfactual rotation of the full layout around a fixed camera. Red placeholders mark variable content.
Figure 9: Representative endpoint-image generation failures. From left to right: object truncation, incorrect object identity, layered composition, object deformation, and visual artifacts.
Figure 10: Training data composition.
Cam. Dir.
Obj. Dir.
World Dir.
Dist.
Occ.
Att.
Avg.
Move Cam. Dir.
Move Obj. Dir.
Rotate Obj. Dir.
Move World Dir.
Move Dist.
Move Occ.
Remove Occ.
Avg.
Rotate Cam. Dir.
Rotate Obj. Dir.
Rotate World Dir.
Avg.
Model
Overall
L1: Static
L2: Local
L3: Global
Qwen3-VL-4B
34.0
49.2
6.0
35.4
46.0
60.6
42.3
39.9
36.9
28.5
24.2
29.5
31.9
51.8
32.9
33.7
31.9
6.3
30.1
22.8
+ Mixed SFT
64.8
84.8
90.8
65.0
62.9
80.5
74.2
76.4
62.8
41.8
43.4
37.9
44.6
48.9
80.1
51.3
77.6
80.5
61.3
73.1
+ Staged SFT
65.7
88.6
88.8
61.2
68.0
80.5
73.2
76.7
70.6
40.2
43.0
42.4
44.6
54.0
83.3
54.0
79.1
79.7
53.8
70.8
+ GRPO
37.7
53.4
7.6
38.8
29.4
63.3
49.5
40.3
14.7
12.5
21.1
33.7
33.3
46.8
62.0
32.0
42.6
63.7
30.8
45.7
Appendix
Table 9: Per-task accuracy (%) of Qwen3-VL-4B and its trained variants on SpaceCast-Bench.
Figure 11: L1 Camera Direction. Rather than recovering the ground-plane layout and expressing it in the first main view’s camera frame, models resort to image-space left–right positions or anchor on the last input view. Correct answers are shown in green; erroneous reasoning and predictions are highlighted in red.
Figure 12: L2 Remove Occlusion. Models treat the removed tap as the sole occluder and conclude that the hand-soap bottle is now fully visible, failing to recompute visibility against the complete remaining scene. The case illustrates a broader pattern in which partial visibility is equated with the absence of occlusion. Correct answers are shown in green; erroneous reasoning and predictions are highlighted in red.
Figure 13: L2 Move World Direction. Models conflate the table’s movement direction with the resulting direction of the mouse relative to the bottle, or substitute coarse room semantics for the required world-coordinate relation. The case further requires propagating the table’s displacement to its rigidly attached mouse, an attachment dependency that models consistently overlook. Correct answers are shown in green; erroneous reasoning and predictions are highlighted in red.
Figure 14: L3 Rotate Camera Direction. Models apply a superficial left–right reversal or reason from an unreformed input view rather than rotating the full scene layout and recomputing the queried relation in the fixed camera frame. The case illustrates how transformation reasoning degrades when it requires maintaining and manipulating a global scene representation. Correct answers are shown in green; erroneous reasoning and predictions are highlighted in red.
Vision language models (VLMs) excel at many tasks but still struggle with spatial reasoning when critical information is not directly observable. Many such problems require imaginative perception: inferring what would be seen from an unseen viewpoint, tracing paths through occluded spaces, or integrating partial observations into a coherent spatial representation. We introduce Imaginative Perception Tokens (IPT), intermediate perceptual representations that externalize what a VLM would perceive under alternative spatial configurations while remaining consistent with the observed input. To study this capability, we formulate three tasks, Perspective Taking (PET), Path Tracing (PT), and Multiview Counting (MVC), and construct datasets of approximately 20K examples with ground truth imaginations, answers, and evaluation benchmarks. Using the unified VLM BAGEL as the backbone, IPT supervision consistently improves spatial reasoning and often outperforms textual chain of thought training, even without generating images at inference time. On MVC, IPT improves accuracy by 3.4% and achieves competitive performance with strong closed-source models on PT. We further find that combining IPT and label-only supervision yields additional gains, whereas textual chain of thought can substantially degrade performance, suggesting a modality mismatch when spatial computation is forced through language. Overall, IPT provides a principled supervision signal for reasoning about unobserved spatial structure, improving generalization while producing interpretable intermediate representations.
Mahtab Bigverdi, Linjie Li, Weikai Huang +8
University of Washington · 2Allen Institute for AI · 3Microsoft +1
Multimodal large language models (MLLMs) excel at visual interpretation but fail on spatial reasoning tasks that humans solve reliably. Existing benchmarks evaluate these models as black boxes, limiting their ability to identify the underlying causes of lower performance: when a model fails a spatial reasoning task, it remains difficult to ascertain whether the hurdle is perceptual, such as recognizing object boundaries, or cognitive, such as reasoning about occlusion to infer hidden geometry. We introduce Spatial-IQ, a hierarchical diagnostic framework that decomposes object counting in stacked 3D structures into 9 perceptual and cognitive sub-tasks organized by the developmental stages of human spatial cognition, with mental rotation as an additional target probe. Using NVIDIA Isaac Sim, we procedurally generated a diverse dataset of roughly 80,000 stacked 3D structures with per-task ground truth. We evaluate models across three output formats (free-response text, multiple-choice images, and image editing) alongside a human baseline. The Spatial-IQ framework shows that top-performing models often succeed at the target task (object counting) without succeeding on the lower-level sub-tasks intended to support it, and that models differ in how much of these hierarchical chains they preserve, often revealing shortcut behavior that raw target-task accuracy alone would obscure. Finally, we demonstrate that training models with chain-of-thought (CoT) supervision over our hierarchical sub-tasks, combined with reinforcement learning with verifiable rewards, significantly improves both spatial consistency across sub-tasks and target-task accuracy, supporting the value of the proposed decomposition as both a diagnostic tool and a training signal.
Multimodal large language models (MLLMs) have achieved remarkable progress in vision-language tasks, but continue to struggle with spatial reasoning. Existing spatial MLLMs rely on large-scale datasets, explicit 3D inputs, architecture-specific modifications, or sparse Reinforcement Learning (RL) methods that provide insufficient guidance for spatially-grounded reasoning. We introduce SpatialThinker. To our knowledge, it is the first MLLM unifying Scene Graph Generation (SGG) and visual reasoning in a single pass via online RL. The model simulates human-like spatial perception by constructing a mental scene graph of task-relevant objects and relations, and reasoning toward an answer via dense spatial rewards. Our contributions are threefold: (1) SGG-grounded reasoning: integrating SGG directly within the reasoning chain rather than as a disjoint preprocessing step; (2) STVQA-7K: a high-quality spatial VQA training dataset via a scalable synthesis pipeline; and (3) a dense spatial reward design that enforces structured grounding during RL and generalizes to improve broad visual perception. SpatialThinker-7B achieves 3.6× larger gains over SFT and 1.7× better in- and out-of-distribution generalization than sparse RL. Trained on only 7K samples, SpatialThinker-7B matches GPT-5 and outperforms GPT-4o, while SpatialThinker-30B surpasses both GPT-5 and Claude 4 Sonnet on average across 14 spatial and real-world benchmarks, demonstrating that structured spatial grounding with reward-aligned reasoning enables robust spatial understanding with limited data.
Hunar Batra, Haoqin Tu, Hardy Chen +3
University of Oxford · University of California, Santa Cruz