Spatial--temporal benchmarks are valid only when their questions, source annotations, and answer options identify the same physical quantity. We audit STI-Bench against the official ScanNet, Waymo, and Omni6DPose sources and find systematic coordinate-system and timestamp errors, under-specified targets and times, and disagreements between keyed options and answer details. We introduce ReSTI, a source-backed revision that reconstructs every recoverable answer under an explicit target, time, coordinate system, physical quantity, and unit. Source reconstruction reveals task-level geometric failures: ScanNet Grounding omits the required alignment between annotation and raw camera coordinate systems, while Orientation measures camera rotation on the wrong plane. ReSTI replaces these labels with explicit, source-consistent geometric definitions and corrects other source-verifiable defects, including Waymo poses evaluated at the wrong timestamp. Across 2,064 legacy questions, ReSTI retains 1,782 questions and records 282 evidence-backed exclusions. ReSTI therefore provides a conservative and source-traceable basis for evaluating precise video spatial--temporal reasoning. Project page: https://github.com/pengzhansun/ReSTI.
Figures & tables
Figure 1 : A valid benchmark item requires three mutually consistent layers. The question contract must specify a unique physical quantity; the source contract must reconstruct that quantity in the correct coordinate and time frames; and the answer interface must encode it without contradictory labels or option-only shortcuts. (A) The pose question omits its target time. (B) The released Grounding computation omits the conversion from annotation coordinates to the raw scan frame expected by the camera pose. (C) In a separate ScanNet example (scene0117_00), the released detail is 8.03∘ and option B is 8∘ , but the key selects option C ( 0∘ ).
Figure 2 : The ReSTI audit and repair pipeline. (1) Specify the question contract and identify the source annotations. (2) Resolve the entity and time, align coordinate frames, and apply the physical operator to reconstruct the target. (3) Compare the released answer with the reconstructed target under a comparable contract and record a retained, corrected, re-contracted, or excluded verdict. (4) Synthesize candidates, validate their structure and geometry, and test for option-only shortcuts. Exclusions carry reasons and evidence; option-only tests guide candidate refinement.
Figure 3 : Audit overview. (a) Accepted and excluded questions by subtask, with exact counts. ReSTI accepts 1,782 of 2,064 questions; accepted includes retained, corrected, and re-contracted items. (b) Distribution of 2,800 defect reports across six categories. Percentages use the sum of these reports, not the number of questions: one question can contribute to multiple categories.
Figure 4 : A Grounding label displaced from its target by 9.15 m. (a) Original ScanNet RGB frame with the source table box projected using its mesh annotation, axis alignment, camera pose, and intrinsics. (b) Top-down view of the source mesh. Green marks the table; the red cross is the signed Answer Detail center mapped back into the raw scene. The dashed line connects their centers; the distance is measured in 3D. The lower panels contrast the invalid frame substitution with the correct transform. Their reported tuples are camera-frame centers; panel (b) displays both centers after transformation into the common scene frame. This displacement is independent of how the legacy scalar heading is interpreted.
Figure 5 : A vertical-plane angle is not horizontal heading change. (a) Original RGB endpoints of ScanNet scene0012_00; times follow the STI video frame indexing. (b) Normalized camera-forward vectors projected onto raw-world X–Z, a vertical plane, reproduce the released detail. (c) Projection onto the gravity-horizontal X–Y plane gives the repaired heading. Blue and gold denote the start and end directions; arcs show signed angles from the start direction to the end direction. The vertical-plane angle includes up/down tilt, whereas the horizontal projection measures heading change on the floor plane. Both calculations use the same original source poses.