Spatial reasoning benchmarks evaluate vision-language models across diverse tasks, but task-level scores do not reveal which underlying capabilities account for success or failure. Each task requires recovering spatial evidence, representing geometry, and reasoning over it. We disentangle these capabilities by comparing predicted and ground-truth spatial context under a shared schema and coordinate contract. This comparison reveals four recurring sources of error: inaccurate perception, missing information in the spatial context, selection of the wrong measurement, and errors in reference frames or in tracking position and orientation. Guided by this diagnosis, we develop CROSS, a training-free library of typed geometric operators and spatial skills that function over available evidence to support reliable video spatial reasoning. The resulting library supplies verified context to non-coding VLMs or callable skills to a SpatialClaw agent. We evaluate \methodname{} on five benchmarks. \methodname{} raises the average score from 55.9% to 60.2% on ReVSI and improves the SpatialClaw result from 62.8% to 66.3% on DSI-Bench. These gains demonstrate that explicit handling of spatial conventions can repair systematic reasoning failures without additional training.
Figures & tables
Figure 1: Illustration of spatial context. For each video, either ground-truth or tool-derived spatial clues populate the coordinate contract, scene geometry, entity observations, and viewer trajectory.
Figure 2: Share of GT-context failures matching the diagnosed mode (dark); pale segments show other failures. Rules and counts: Appendix A.5 .
Figure 3: Overview of CROSS . Overview of CROSS. Operators derived from diagnosed failures (left) are composed into task-specific skills (middle) that serve two interfaces (right): context-augmented VLMs receive gated and filtered skill outputs or fall back to video-only reasoning, and code-writing agents call the skills and operators directly.
Family
Role
Frame
Align bases and poses.
Geometry/ shape
Measure surfaces, extents, and floor area.
Set/decision
Deduplicate, count, and match with margins.
Temporal/ state
Integrate paths; update position and heading.
Table 1: Atomic operator families.
Video
Predicted context
GT context
Code writing
Subtask
n
CoT
Base
+ CROSS
Base
+ CROSS
SpatialClaw
+ CROSS
ReVSI-Tiny
Object counting
118
58.2
46.2
58.2 (+12.0)
93.1
94.8 (+1.7)
42.4
45.4 (+3.0)
Absolute distance
239
54.5
48.1
63.5 (+15.4)
58.4
100.0 (+41.6)
44.8
49.6 (+4.8)
Object size
215
67.9
54.6
67.9 (+13.3)
100.0
100.0
37.1
37.2 (+0.1)
Room size
30
55.0
62.0
55.0
56.7
56.0
31.3
33.3 (+2.0)
Table 2: Diagnosis and CROSS gains on ReVSI-Tiny and ReSTI-Tiny . Both context and code-writing arms use Qwen3-VL-32B-Thinking. Green parentheses show gains over the paired baseline.
Figure 4: Relative-direction correction through code-writing (left pair) and context augmentation (right pair): both baselines answer right, whereas both CROSS runs answer left.
System
Frames
Numerical questions (MRA)
Multiple-choice questions (Acc.)
Avg. ↑
Obj. cnt.
Abs. dist.
Obj. size
Room size
Rel. dist.
Rel. dir.
Route plan
Baselines
Chance (random)
All
–
–
–
–
23.7
26.8
26.0
–
Chance (frequency)
All
52.2
40.1
17.4
20.9
25.8
31.9
30.2
31.4
Commercial models
GPT-5.2
64
56.2
41.5
73.9
63.0
48.4
34.9
38.2
50.9
Table 3: Detailed comparison on ReVSI. Numerical questions use MRA and multiple-choice questions use accuracy. Green parentheses show positive score-point gains over the corresponding SpatialClaw or direct-video baseline.
32 key frames shown to the model; up to 64 frames for reconstruction
Perception
SAM3 and DA3 in streaming mode (predicted context; Appendix B )
SAM3 and DA3 in batch mode, called as agent tools
Agent loop
–
Up to 30 code steps with a 600 s limit per step; planning enabled
Appendix
Table 7: Inference settings for the two interfaces.
Subtask
Matching rule
Matched/errors
Wrong measurement
Absolute distance
Prediction within 5% of center-to-center distance
225/227
Wrong frame or state
Relative direction
Left–right mirror (14), reversed facing (5), or fixed scene axes (4)
23/25
Route planning
Incorrect answer differs from the target only by left–right turn swaps
21/40
Pose estimation
Nearest option to ego-as-world (22), unrotated displacement (6), or initial pose (4)
32/48
Appendix
Table 8: Rules and denominators for Figure 2 . Counts are matching errors over all errors in the same GT-context Base run as Table 2 ; percentages are conditional on failure. Multiple rules within a row use first-match assignment.
Bind motion to an interval, distinguish endpoint from path quantities, classify local turns, and persistently update route position and heading.
Appendix
Table 9: Full atomic operator inventory and family responsibilities. The shared safeguard g controls whether derived evidence enters the augmented context and is separate from these spatial computations.
Skill
Fixed operator graph
Main fallback condition
Counting
deduplicate → cardinality
Missing/merged objects or severe track fragmentation.
Only a global envelope or incomplete floor is available.
Appendix
Table 10: Subtask skills and principal validity failures. Each operator sequence runs over the full structured spatial context and is followed by gating and filtering in the context-augmented realization.
Setting
Numerical questions (MRA)
Multiple-choice questions (Acc.)
Average ↑
Obj. cnt.
Abs. dist.
Obj. size
Room size
Rel. dist.
Rel. dir.
Route plan
App. order
Direct
60.8
50.2
73.9
61.3
57.2
57.5
48.7
68.7
59.8
32-frame context
40.6
39.4
45.4
65.5
56.5
53.1
42.4
68.9
51.5
64-frame context
41.9
39.0
46.0
64.7
56.8
53.4
38.7
69.2
51.2
Screened context
42.3
41.2
46.9
57.0
55.9
62.6
46.6
69.7
52.8
CROSS
60.8
53.4
73.9
65.5
57.2
62.6
49.2
68.7
61.4
Appendix
Table 11: Ablations on VSI-Bench with 64 video frames. Average gives equal weight to eight subtasks. Numerical tasks use MRA and categorical tasks use accuracy (higher is better). Bold marks the best result in each column, including ties.
Figure 7: GT-context failures on absolute distance and room size (top) and path length and pose estimation (bottom): center distance used for the surface gap, the scene envelope used for floor area, a sparse trajectory that underestimates path length, and an ego displacement not rotated into the world frame.
Figure 8: GT-context failures on relative direction and route planning (top) and 3D grounding and dimensional measurement (bottom): a handedness reversal, a heading that is never updated, a referring expression that matches many entities, and answer options closer than the context’s precision.
Figure 9: Pose-estimation correction through context augmentation (top) and SpatialClaw (bottom). Both baselines choose A; both CROSS runs compose the camera motion with the initial pose and choose B.
Figure 10: Absolute-distance correction. The baseline uses center distance and returns 3.4 m; CROSS derives the closest-surface gap and returns 2.17 m against a 2.1 m reference. The gate retains the ambiguous couch-role warning while reducing candidate pairs by minimum gap.
Figure 11: Route-planning correction. The baseline selects A (left, right); CROSS supplies the sequence computed with heading-state updates and the model selects C (right, left). Multiple bed and chair tracks remain in the bound roles.
Figure 12: Object-counting recovery through video fallback. The baseline counts four predicted pink-pillow tracks; the deployed gate routes to video and the model returns the reference count of two. Forcing the derived count still returns four, distinguishing the routing benefit from instance reduction.
School of Computer Science, Peking University, Beijing, China · School of Intelligent Science and Technology, Nanjing University, Nanjing, China · Beijing Academy of Artificial Intelligence, Beijing, China