From Reasoning Failures to Composable Video Spatial Intelligence
Organizations: National University of Singapore · University of Science and Technology of China
Abstract
Spatial reasoning benchmarks evaluate vision-language models across diverse tasks, but task-level scores do not reveal which underlying capabilities account for success or failure. Each task requires recovering spatial evidence, representing geometry, and reasoning over it. We disentangle these capabilities by comparing predicted and ground-truth spatial context under a shared schema and coordinate contract. This comparison reveals four recurring sources of error: inaccurate perception, missing information in the spatial context, selection of the wrong measurement, and errors in reference frames or in tracking position and orientation. Guided by this diagnosis, we develop CROSS, a training-free library of typed geometric operators and spatial skills that function over available evidence to support reliable video spatial reasoning. The resulting library supplies verified context to non-coding VLMs or callable skills to a SpatialClaw agent. We evaluate \methodname{} on five benchmarks. \methodname{} raises the average score from 55.9% to 60.2% on ReVSI and improves the SpatialClaw result from 62.8% to 66.3% on DSI-Bench. These gains demonstrate that explicit handling of spatial conventions can repair systematic reasoning failures without additional training.
Figures & tables
| Family | Role |
| Frame | Align bases and poses. |
| Geometry/ shape | Measure surfaces, extents, and floor area. |
| Set/decision | Deduplicate, count, and match with margins. |
| Temporal/ state | Integrate paths; update position and heading. |
| Video | Predicted context | GT context | Code writing | |||||
| Subtask | CoT | Base | CROSS | Base | CROSS | SpatialClaw | CROSS | |
| ReVSI-Tiny | ||||||||
| Object counting | 118 | 58.2 | 46.2 | 58.2 (+12.0) | 93.1 | 94.8 (+1.7) | 42.4 | 45.4 (+3.0) |
| Absolute distance | 239 | 54.5 | 48.1 | 63.5 (+15.4) | 58.4 | 100.0 (+41.6) | 44.8 | 49.6 (+4.8) |
| Object size | 215 | 67.9 | 54.6 | 67.9 (+13.3) | 100.0 | 100.0 | 37.1 | 37.2 (+0.1) |
| Room size | 30 | 55.0 | 62.0 | 55.0 | 56.7 | 56.0 | 31.3 | 33.3 (+2.0) |
| System | Frames | Numerical questions (MRA) | Multiple-choice questions (Acc.) | Avg. | |||||
| Obj. cnt. | Abs. dist. | Obj. size | Room size | Rel. dist. | Rel. dir. | Route plan | |||
| Baselines | |||||||||
| Chance (random) | All | – | – | – | – | 23.7 | 26.8 | 26.0 | – |
| Chance (frequency) | All | 52.2 | 40.1 | 17.4 | 20.9 | 25.8 | 31.9 | 30.2 | 31.4 |
| Commercial models | |||||||||
| GPT-5.2 | 64 | 56.2 | 41.5 | 73.9 | 63.0 | 48.4 | 34.9 | 38.2 | 50.9 |
Appendix figures & tables11 assets
Supplementary material from the paper’s appendix.
Appendix
| Context augmentation | Code writing (SpatialClaw) | |
| Model | Qwen3-VL-32B-Thinking | Qwen3-VL-32B-Thinking ( ReVSI-Tiny , ReSTI-Tiny ); Gemma-4-31B-FP8 (ReVSI, DSI-Bench) |
| Serving | vLLM, bf16 | vLLM |
| Decoding | Sampling with temperature 1.0, top- 0.95, top- 20; up to 40,960 new tokens | Temperature 0.6; up to 32,768 (Qwen) or 131,072 (Gemma) tokens per call |
| Video | Uniform frames: 32 ( ReVSI-Tiny ), 30 ( ReSTI-Tiny , STI-Bench), 64 (ReVSI, VSI-Bench) | 32 key frames shown to the model; up to 64 frames for reconstruction |
| Perception | SAM3 and DA3 in streaming mode (predicted context; Appendix B ) | SAM3 and DA3 in batch mode, called as agent tools |
| Agent loop | – | Up to 30 code steps with a 600 s limit per step; planning enabled |
| Subtask | Matching rule | Matched/errors |
| Wrong measurement | ||
| Absolute distance | Prediction within 5% of center-to-center distance | 225/227 |
| Wrong frame or state | ||
| Relative direction | Left–right mirror (14), reversed facing (5), or fixed scene axes (4) | 23/25 |
| Route planning | Incorrect answer differs from the target only by left–right turn swaps | 21/40 |
| Pose estimation | Nearest option to ego-as-world (22), unrotated displacement (6), or initial pose (4) | 32/48 |
| Family | Operators | Responsibility |
| Frame | relative_vector , build_ego_basis , project_local , relative_se3 , compose_se3 , pose_distance | Construct directed vectors and a handedness-aware forward/right/up basis, transport poses between declared frames, and score pose candidates. |
| Geometry/shape | clean_geometry , aggregate_geometry , surface_distance , oriented_extent , floor_polygon , polygon_area | Robustly aggregate predicted evidence and compute the benchmark-defined surface, extent, or footprint quantity. |
| Set/decision | deduplicate , cardinality , reduce_set , rank_with_margin , match_candidate | Operate over instance sets and match a computed result to candidates only when its margin exceeds estimated error. |
| Temporal/state | endpoint_displacement , path_integral , interval_rate , classify_sector , advance_agent | Bind motion to an interval, distinguish endpoint from path quantities, classify local turns, and persistently update route position and heading. |
| Skill | Fixed operator graph | Main fallback condition |
| Counting | deduplicate cardinality | Missing/merged objects or severe track fragmentation. |
| Absolute distance | select/clean/aggregate surface_distance reduce(min) | A role is absent or geometric uncertainty is high. |
| Relative distance | Absolute-distance graph per candidate, then reduce(min) rank_with_margin | Candidate coverage is incomplete or the rank margin is too small. |
| Relative direction | relative_vector basis/project/classify | Missing role, degenerate facing, or boundary ambiguity. |
| Object size | select/clean/aggregate oriented_extent reduce(max) | Truncation or unstable multi-view extent. |
| Room area | select_region clean floor_polygon polygon_area | Only a global envelope or incomplete floor is available. |
| Setting | Numerical questions (MRA) | Multiple-choice questions (Acc.) | Average | ||||||
| Obj. cnt. | Abs. dist. | Obj. size | Room size | Rel. dist. | Rel. dir. | Route plan | App. order | ||
| Direct | 60.8 | 50.2 | 73.9 | 61.3 | 57.2 | 57.5 | 48.7 | 68.7 | 59.8 |
| 32-frame context | 40.6 | 39.4 | 45.4 | 65.5 | 56.5 | 53.1 | 42.4 | 68.9 | 51.5 |
| 64-frame context | 41.9 | 39.0 | 46.0 | 64.7 | 56.8 | 53.4 | 38.7 | 69.2 | 51.2 |
| Screened context | 42.3 | 41.2 | 46.9 | 57.0 | 55.9 | 62.6 | 46.6 | 69.7 | 52.8 |
| CROSS | 60.8 | 53.4 | 73.9 | 65.5 | 57.2 | 62.6 | 49.2 | 68.7 | 61.4 |