Embodied agents must reason about 3D space while the video is still arriving, answering questions as soon as they have observed enough of the scene. VLMs that incorporate 3D geometric priors achieve strong spatial reasoning, but they operate offline, i.e., the full video must be available before they produce an answer. Streaming VLMs process frames causally and decide for themselves when to respond, yet they lack explicit 3D representations. We present SpaTime, a streaming VLM that fuses causal geometry tokens into the language model at every frame, using only the frames observed so far. To supervise when the model answers, we propose a response-time loss that maps per-frame response probabilities to a differentiable expected response time and penalizes the distance from the ground-truth frame. For evaluation, we construct StreamVSTI-Bench and StreamVSI-Bench, streaming adaptations of VSTI-Bench and VSI-Bench. On StreamVSTI-Bench, SpaTime reaches 49.2% overall accuracy and reduces the mean response-time error by 66% relative to the strongest streaming baseline.
Figures & tables
Figure 1: Offline VLMs lack real-time answering capability, and general streaming VLMs lack 3D awareness. SpaTime fuses visual and causal 3D geometry tokens to achieve a Streaming VLM for spatial reasoning.
Figure 2: SpaTime pipeline. Figure (a) illustrates the overall pipeline. Each frame is encoded by a frozen visual encoder and a frozen causal geometry encoder, then (b) a lightweight projector aligns the geometry tokens with the visual tokens, and fused via element-wise addition. In (c), a response time loss Ltime aligns the predicted response time with the ground truth.
Figure 3: Data curation. For each question, a geometry-driven visibility check and a VLM check locate the frames where its objects are visible. Next, a per-type rule selects the answer frame, and a query frame is assigned to create the streaming data.
Table 4
Ablations
Metrics
Geo. Tokens
RT Loss
Strict
Charitable
Δt (s) ↓
43.99
44.05
0.35
✓
47.93
48.37
0.18
✓
✓
48.80
49.20
0.12
Table 3: Ablation on StreamVSTI-Bench. Components are added cumulatively to a fine-tuned Streamo baseline; the shaded row is our final model.
Method
StreamVSTI
StreamVSI
VideoLLM-online
6.22
4.94
Dispider
23.39
0.63
Streamo
43.99
21.09
Ours
48.80
29.44
Table 4: Strict-mode average accuracy (%) for streaming methods. The strict protocol scores only the earliest predicted ⟨Response⟩ and counts a missing response as zero, penalizing timing failures directly. Offline models do not respond and are omitted.
Ablations
Δt (s) ↓
Geo. Tokens
RT Loss
All
Future-only
✓
0.18
0.52
✓
✓
0.12
0.35
Table 5: The response time loss targets future questions. Mean response-time error Δt (s, lower is better) over all questions versus current-ask-future questions only, whose answers depend on frames that arrive after the query. Adding the response time loss (row 2) cuts the error most on future questions, where the model must decide how long to wait.
Figure A1: Interface for the human verification study of the visibility annotations. For each of the 500 sampled object–frame pairs, annotators see the projected 3D box overlay, a zoomed crop, and neighboring frames around the pipeline’s first-seen claim, and judge whether the object is genuinely recognizable; the resulting human–pipeline agreement is 84.4% .
Appendix figures & tables4 assets
Supplementary material from the paper’s appendix.
Appendix
Hyperparameter
Value
LoRA rank / α / dropout
128 / 256 / 0.05
LoRA targets
q,k,v,o,gate,up,down
Fully trained modules
embed_tokens , lm_head
Trainable parameters
1,461.2 M ( 15.0% )
Base / projector learning rate
1×10−5 / 5×10−4
Optimizer
AdamW, β=(0.9,0.95)
Appendix
Table A1: Training hyperparameters of SpaTime.
Split
Future
Current
Past
Train ( 132,568 )
41,196
22,268
69,104
Test ( 6,042 ) ∗
2,076
2,075
1,891
Appendix
Table A2: Temporal-channel distribution of StreamVSTI-Bench.
Channel
Bucket
n
Acc. (%)
Exact round
Δ round
Current
MC
1,440
62.50
99.1%
+0.001
Current
Num.
635
35.04
98.7%
+0.005
Future
MC
1,512
59.92
60.4%
−0.371
Future
Num.
564
35.32
74.3%
−0.242
Past
MC
1,346
63.45
100.0%
0
Past
Num.
545
36.90
99.8%
0
Appendix
Table A3: Per-temporal-channel breakdown on StreamVSTI-Bench. “Exact round” is the fraction of samples whose ⟨Response⟩ lands on precisely the ground-truth round; Δ round is the mean of tpred−tgt , so negative values mean the model answers early.
Task / channel
Scene
Question (abridged)
GT (round, ans.)
Pred. (round, ans.)
Rel. pos. (lr) / future
scene0700_02
At 33.40s, will telephone be to the [Left/Right] relative to keyboard? (asked at round 14)
(34, A)
(34, A)
Rel. dist. (v3) / future
scene0580_00
Measuring from the closest point of each object at 43.42s, which of (nightstand, bed) will be closest to the camera? (asked at round 24)
(44, B)
(44, B)
Displacement / future
scene0664_00
How far (in meters) will the camera move between 13.12s and 39.35s? (asked at round 14)
(39, 0.5)
(39, 0.5)
Rel. pos. (ud) / current
scene0231_02
At current time, relative to backpack, is window to the [Up/Down]?
(4, A)
(4, A)
Rel. dist. (v2) / current
scene0580_01
Measuring from the closest point of each object at current time, which of (backpack, table, bed) is closest to the camera?
(39, C)
(39, C)
Movement dir. / past
scene0050_00
Looking back, what was the primary consistent direction of the camera’s movement from 12.11s to 84.80s?
(109, C)
(109, C)
Appendix
Table A4: Qualitative examples across task families and temporal channels on StreamVSTI-Bench. For every example, the predicted ⟨Response⟩ round matches the ground-truth round exactly and the answer is correct; future-channel examples require holding ⟨Standby⟩ for 20 – 25 rounds before responding.