Spatial--temporal benchmarks are valid only when their questions, source annotations, and answer options identify the same physical quantity. We audit STI-Bench against the official ScanNet, Waymo, and Omni6DPose sources and find systematic coordinate-system and timestamp errors, under-specified targets and times, and disagreements between keyed options and answer details. We introduce ReSTI, a source-backed revision that reconstructs every recoverable answer under an explicit target, time, coordinate system, physical quantity, and unit. Source reconstruction reveals task-level geometric failures: ScanNet Grounding omits the required alignment between annotation and raw camera coordinate systems, while Orientation measures camera rotation on the wrong plane. ReSTI replaces these labels with explicit, source-consistent geometric definitions and corrects other source-verifiable defects, including Waymo poses evaluated at the wrong timestamp. Across 2,064 legacy questions, ReSTI retains 1,782 questions and records 282 evidence-backed exclusions. ReSTI therefore provides a conservative and source-traceable basis for evaluating precise video spatial--temporal reasoning. Project page: https://github.com/pengzhansun/ReSTI.
Figures & tables
Figure 1 : A valid benchmark item requires three mutually consistent layers. The question contract must specify a unique physical quantity; the source contract must reconstruct that quantity in the correct coordinate and time frames; and the answer interface must encode it without contradictory labels or option-only shortcuts. (A) The pose question omits its target time. (B) The released Grounding computation omits the conversion from annotation coordinates to the raw scan frame expected by the camera pose. (C) In a separate ScanNet example (scene0117_00), the released detail is 8.03∘ and option B is 8∘ , but the key selects option C ( 0∘ ).
Figure 2 : The ReSTI audit and repair pipeline. (1) Specify the question contract and identify the source annotations. (2) Resolve the entity and time, align coordinate frames, and apply the physical operator to reconstruct the target. (3) Compare the released answer with the reconstructed target under a comparable contract and record a retained, corrected, re-contracted, or excluded verdict. (4) Synthesize candidates, validate their structure and geometry, and test for option-only shortcuts. Exclusions carry reasons and evidence; option-only tests guide candidate refinement.
Figure 3 : Audit overview. (a) Accepted and excluded questions by subtask, with exact counts. ReSTI accepts 1,782 of 2,064 questions; accepted includes retained, corrected, and re-contracted items. (b) Distribution of 2,800 defect reports across six categories. Percentages use the sum of these reports, not the number of questions: one question can contribute to multiple categories.
Figure 4 : A Grounding label displaced from its target by 9.15 m. (a) Original ScanNet RGB frame with the source table box projected using its mesh annotation, axis alignment, camera pose, and intrinsics. (b) Top-down view of the source mesh. Green marks the table; the red cross is the signed Answer Detail center mapped back into the raw scene. The dashed line connects their centers; the distance is measured in 3D. The lower panels contrast the invalid frame substitution with the correct transform. Their reported tuples are camera-frame centers; panel (b) displays both centers after transformation into the common scene frame. This displacement is independent of how the legacy scalar heading is interpreted.
Figure 5 : A vertical-plane angle is not horizontal heading change. (a) Original RGB endpoints of ScanNet scene0012_00; times follow the STI video frame indexing. (b) Normalized camera-forward vectors projected onto raw-world X–Z, a vertical plane, reproduce the released detail. (c) Projection onto the gravity-horizontal X–Y plane gives the repaired heading. Blue and gold denote the start and end directions; arcs show signed angles from the start direction to the end direction. The vertical-plane angle includes up/down tilt, whereas the horizontal projection measures heading change on the floor plane. Both calculations use the same original source poses.
Spatial reasoning from egocentric videos is inherently challenging because the observable evidence is constrained by the camera trajectory. Existing methods rely on single-turn inference, forcing models to resolve geometric ambiguity through semantic priors rather than verifiable evidence. We argue that spatial reasoning should be revisitable: conclusions formed under limited evidence should remain open to revision when complementary viewpoints become available. Building on this insight, we propose Reason, then Re-reason (ReRe), a training-free, inference-time framework with two phases: in the Reason Phase, an MLLM forms a spatial hypothesis from the original video; in the Re-reason Phase, it verifies or revises the hypothesis by observing a synthesized novel-view video. To enable effective cross-view revisiting, we design a Geometry-to-Video pipeline that renders strategically complementary novel views from predicted 3D geometry. These views feature an elevated, oblique perspective with scene-spanning coverage, while preserving the MLLM's native video interface without architectural modifications. Extensive evaluations on VSI-Bench and STI-Bench demonstrate that ReRe substantially boosts open-source MLLMs to rival proprietary state-of-the-art performance. Project page: https://zhenjiemao.github.io/ReRe/
Chaofan Ma, Zhenjie Mao, Yuhuan Yang +5
Cooperative Medianet Innovation Center, Shanghai Jiao Tong University · Shanghai Jiao Tong University · Tongji University
Spatial reasoning benchmarks evaluate vision-language models across diverse tasks, but task-level scores do not reveal which underlying capabilities account for success or failure. Each task requires recovering spatial evidence, representing geometry, and reasoning over it. We disentangle these capabilities by comparing predicted and ground-truth spatial context under a shared schema and coordinate contract. This comparison reveals four recurring sources of error: inaccurate perception, missing information in the spatial context, selection of the wrong measurement, and errors in reference frames or in tracking position and orientation. Guided by this diagnosis, we develop CROSS, a training-free library of typed geometric operators and spatial skills that function over available evidence to support reliable video spatial reasoning. The resulting library supplies verified context to non-coding VLMs or callable skills to a SpatialClaw agent. We evaluate \methodname{} on five benchmarks. \methodname{} raises the average score from 55.9% to 60.2% on ReVSI and improves the SpatialClaw result from 62.8% to 66.3% on DSI-Bench. These gains demonstrate that explicit handling of spatial conventions can repair systematic reasoning failures without additional training.
Current evaluations of spatial intelligence can be systematically invalid under modern vision-language model (VLM) settings. First, many benchmarks derive question-answer (QA) pairs from point-cloud-based 3D annotations originally curated for traditional 3D perception. When such annotations are treated as ground truth for video-based evaluation, reconstruction and annotation artifacts can miss objects that are clearly visible in the video, mislabel object identities, or corrupt geometry-dependent answers (e.g., size), yielding incorrect or ambiguous QA pairs. Second, evaluations often assume full-scene access, while many VLMs operate on sparsely sampled frames (e.g., 16-64), making many questions effectively unanswerable under the actual model inputs. We improve evaluation validity by introducing ReVSI, a benchmark and protocol that ensures each QA pair is answerable and correct under the model's actual inputs. To this end, we re-annotate objects and geometry across 381 scenes from 5 datasets to improve data quality, and regenerate all QA pairs with rigorous bias mitigation and human verification using professional 3D annotation tools. We further enhance evaluation controllability by providing variants across multiple frame budgets (16/32/64/all) and fine-grained object visibility metadata, enabling controlled diagnostic analyses. Evaluations of general and domain-specific VLMs on ReVSI reveal systematic failure modes that are obscured by prior benchmarks, yielding a more reliable and diagnostic assessment of spatial intelligence.
Yiming Zhang, Jiacheng Chen, Jiaqi Tan +3
1Simon Fraser University · 2Hong Kong University of Science and Technology · University of Waterloo +1