Real-world AI systems must reason about objects that are no longer visible: an AR assistant guiding a user back to an object used earlier, a household robot retrieving an item someone put away. This requires not just recalling where an object was last seen, but updating its state when it is moved and retaining that update once it leaves view. We refer to this as out-of-sight spatiotemporal reasoning. We introduce Beyond3D, the first VQA benchmark to isolate this ability in dynamic egocentric video: every query targets an object that has been relocated and has since left the field of view. We create our questions from HD-EPIC annotations, building a visibility track for each dynamic object from its 3D position, the camera pose, and the scene geometry to understand at each moment whether it is visible, occluded, or out of view. Beyond3D comprises 9,000 questions in eight types over 135 videos from nine participants, organized as one reasoning chain: visual grounding (is the target observable now), temporal grounding (when it was last visible and last placed), scene localization (which fixture anchors that location), and 3D spatial perception (where it lies relative to the current viewpoint or another object in the scene). We benchmark nine general-purpose and spatially specialized VLMs. The best model reaches 42.2% against 29.7% chance and text-only baselines reaching 31.9%, with the largest failures in recovering when an object was last visible, showing that tracking object movement out of sight remains far from solved for current VLMs.
Figures & tables
Benchmark
Ego.
3D
Tem- Loc.
Spa- Upd.
Diag.
Upd- OOS
Geo- Vis.
(A) Temporal and object memory
EgoTempo [ 27 ]
✓
–
∘
∘
–
–
–
EgoMemReason [ 36 ]
✓
–
∘
∘
–
–
–
(B) Static 3D spatial reasoning
VSI-Bench [ 39 ]
✓
✓
–
–
–
–
–
(C) Dynamic spatial state reasoning
Table 1 : Comparison with related spatial and memory benchmarks. ✓ : explicitly evaluated; ∘ : partially covered or not systematically enforced; – : not explicitly targeted. Ego. : egocentric input; 3D : explicit 3D spatial reasoning; Tem-Loc. : temporal localization; Spa-Upd. : spatial-state update; Diag. : multi-level diagnostic decomposition; Upd-OOS. : updated target queried after loss of visibility; Geo-Vis. : geometry-aware visibility.
Figure 2 : Beyond3D benchmark construction. (1) Visibility tracks are inferred through view, occlusion, and detection checks. (2) Valid out-of-sight anchors are converted into 9,000 questions.
Figure 3 : Cross-view object projection. An Knife’s annotated image footprint in Camera View 1 is back-projected to its last known 3D depth and reprojected into Camera View 2 using the relative camera pose. The projected footprint approximates the object’s spatial extent under the new viewpoint.
Figure 4 : Evaluation-set statistics. Distribution of the 1,000 out-of-sight query anchors over query time ( a ), out-of-sight horizon ( b ), and number of times the target was moved before the query ( c ), together with the query-time distribution of the 1,000 visible control anchors ( d ). Dashed lines mark the temporal and horizon stratification boundaries used for sampling.
Model
Size
Macro Avg.
Visual Grounding
Temporal Grounding
Scene Localization
3D Spatial Perception
Visibility Check
Last Visible Time
Last Placement Time
Nearest Fixture
Object–Camera Direction
Object–Camera Distance
Object–Object Direction
Object–Object Distance
No. of questions
9000
2000
1000
1000
1000
1000
1000
1000
1000
Random guessing
–
29.7
50.0
20.0
20.0
22.7
25.0
33.3
33.3
33.3
Text Only
General-Purpose Models
Qwen-3.6 [ 30 ]
35B-A3B
31.0
50.0
20.9
21.2
33.4
24.6
32.5
32.6
33.0
Table 2 : VLM accuracy (%) on Beyond3D . Performance of VLMs across question types under Text only , Video + Text and Last Frame Only + Text input settings. Bold entries indicate the best-performing model for the respective question type.
Figure 5 : Performance across temporal conditions. Left: Accuracy across short, medium, and long out-of-sight horizons. Right: Accuracy across early, middle, and late query times.
Figure 6 : Cumulative accuracy over temporal distance. Accuracy when progressively accepting the rounded ground truth (GT) and timestamp choices up to Near ( ± 1–2 s), Medium ( ± 3–4 s), and Far ( ± 5–6 s). Top: full-video models; Bottom: context-limited models (InternVL-3.5 / Cambrian-P / VLM-3R: 150 frames, Spatial-MLLM: 64 frames). Left: last-visible time; Right: last-placement time.
Figure 7 : Visibility-state estimation. Accuracy in determining whether the queried object is visible at query time, reported separately for objects that are not visible and visible .
Figure 8 : Prediction distributions for 3D spatial perception. Stacked bars show prediction frequencies (%). (a) object–camera direction (FL/FR/BL/BR: front/back left/right); (b) object–camera distance; (c) camera-aligned object–object direction; (d) object–object distance.
Figure 9 : Effect of temporal cues on Qwen-3.6-27B. Left: Macro accuracy across out-of-sight (OOS) horizons, with shading indicating the improvement over the baseline. Right: Paired improvement for visibility, nearest-fixture, object–camera (O–C), and object–object (O–O) direction and distance questions. Error bars show 95% bootstrapped confidence intervals.
Figure 10 : Qualitative analysis. For each example, we show the provided temporal cues (blue), keyframes and Q&A, and reasoning traces. Correct and incorrect steps are marked in green and red. Top: the model misses the final relocation of the blueberry box and misidentifies it. Bottom: the model correctly tracks and grounds the food processing lid but fails in camera-relative 3D reasoning.
Question
Exact template
Answer choices
Visual Grounding
Visibility Check
At [TIME] , is the previously moved [OBJECT] visible in the current frame?
No ; Yes
Temporal Grounding
Last Visible Time
Which timestamp is closest to when the [OBJECT] was last visible?
5 timestamps: HH:MM:SS --- N seconds before the end
Last Placement Time
The [OBJECT] was moved earlier in the video. Which timestamp is closest to when it last stopped being moved?
5 timestamps: HH:MM:SS --- N seconds before the end
Scene Localization
Table 3 : Question templates and answer choices. [TIME] , [OBJECT] , and [REF] denote the query timestamp, target object, and visible reference object, respectively.
Figure 11 : Controlled correct-answer distributions. (a–b) Chronological answer-rank distributions for Last Visible Time and Last Placement Time . Chronological rank is obtained by sorting the five displayed timestamp choices from earliest to latest; rank 1 is the earliest and rank 5 the latest. Each rank is correct for exactly 200 of 1,000 questions in each step. (c–f) Semantic correct-answer distributions for the four 3D spatial perception questions, grouped by answer meaning rather than the shuffled multiple-choice option letter. Panel (c) has four classes and is exactly balanced; panels (d–f) have three classes and are balanced to 334/333/333 . FL/FR/BL/BR denote front/back left/right, and distance values are in metres. Dashed lines mark equal-share reference counts.
Figure 12 : Correct-answer distribution for Nearest Fixture . Nearest Fixture asks for the fixture type or counter area closest to the target object’s last known position. Bars show the number of questions for each correct label. Unlike the other question types, this distribution is not balanced but retains the naturally occurring fixture frequencies. All 47 labels occurring across the nine kitchens are shown, and the counts sum to the 1,000 out-of-sight anchors.
Figure 13 : Example preprocessed query frame. The timestamp is placed in the masked region outside the fisheye field of view. For object-relative spatial questions, the visible reference object is indicated by a red marker.
Figure 14 : Object location inference from HD-EPIC annotations. The object is held at its first annotated start location before the first movement, marked in_motion during a movement, and held at the endpoint of its most recent completed movement thereafter.
Figure 15 : Detected-fraction distribution of scored intervals. Share of intervals (%) whose detected fraction of tested frames falls in each 10% bin ( n=445,301 ). The distribution is strongly U-shaped: 65.2% of intervals fall in the lowest bin and 20.1% in the highest, leaving under a sixth of intervals spread across the middle.
Pipeline state
# Marks
# Visible
Acc. (%)
out_of_view
1,293
105
91.9
occluded
291
11
96.2
visually_unconfirmed
1,510
372
75.4
detected_visible
1,008
819
81.2
not_groundable
70
−
−
Table 4 : Human audit of the visibility tracks. We compare the total number of annotated marks per visibility track state with the number of marks judged to be visible. These correspond to disagreements for the first three states, which the pipeline labels not visible, and agreements for detected_visible . Acc. is computed as the percentage of agreeing marks. The pipeline makes no claim for not_groundable . The four scored states sum to the 4,102 marks used for scoring while the 4 in_motion marks among the 4,176 gold marks are omitted, as the pipeline derives no visibility evidence for them.
Metric
All marks
95% CI
Unanimous
Accuracy
83.5
[80.7,85.8]
87.6
Precision
81.2
[76.9,85.2]
82.2
Recall
62.7
[54.9,68.5]
69.3
F1
70.8
–
75.2
Balanced acc.
78.0
[74.1,81.0]
81.9
Audited marks
4,102
518
Table 5 : Pipeline versus majority human vote. Metrics over raw mark counts with 95% cluster-bootstrap confidence intervals over videos. The last column restricts scoring to the 518 scored marks among the 520 multiply-rated marks on which the annotators are unanimous.
Part.
Marks
Acc.
Prec.
Rec.
Bal. acc.
P01
474
84.0
77.7
63.0
77.8
P02
426
81.7
75.3
57.5
74.7
P03
323
80.8
65.9
61.4
74.7
P04
654
87.9
89.4
69.8
83.0
P05
204
81.9
93.6
69.5
82.2
P06
557
82.9
79.9
66.8
79.1
Table 6 : Audit results per participant. Percentages over raw mark counts, with Marks the number of scored marks.