Embodied agents must reason about 3D space while the video is still arriving, answering questions as soon as they have observed enough of the scene. VLMs that incorporate 3D geometric priors achieve strong spatial reasoning, but they operate offline, i.e., the full video must be available before they produce an answer. Streaming VLMs process frames causally and decide for themselves when to respond, yet they lack explicit 3D representations. We present SpaTime, a streaming VLM that fuses causal geometry tokens into the language model at every frame, using only the frames observed so far. To supervise when the model answers, we propose a response-time loss that maps per-frame response probabilities to a differentiable expected response time and penalizes the distance from the ground-truth frame. For evaluation, we construct StreamVSTI-Bench and StreamVSI-Bench, streaming adaptations of VSTI-Bench and VSI-Bench. On StreamVSTI-Bench, SpaTime reaches 49.2% overall accuracy and reduces the mean response-time error by 66% relative to the strongest streaming baseline.
Figures & tables
Figure 1: Offline VLMs lack real-time answering capability, and general streaming VLMs lack 3D awareness. SpaTime fuses visual and causal 3D geometry tokens to achieve a Streaming VLM for spatial reasoning.
Figure 2: SpaTime pipeline. Figure (a) illustrates the overall pipeline. Each frame is encoded by a frozen visual encoder and a frozen causal geometry encoder, then (b) a lightweight projector aligns the geometry tokens with the visual tokens, and fused via element-wise addition. In (c), a response time loss Ltime aligns the predicted response time with the ground truth.
Figure 3: Data curation. For each question, a geometry-driven visibility check and a VLM check locate the frames where its objects are visible. Next, a per-type rule selects the answer frame, and a query frame is assigned to create the streaming data.
Table 4
Ablations
Metrics
Geo. Tokens
RT Loss
Strict
Charitable
Δt (s) ↓
43.99
44.05
0.35
✓
47.93
48.37
0.18
✓
✓
48.80
49.20
0.12
Table 3: Ablation on StreamVSTI-Bench. Components are added cumulatively to a fine-tuned Streamo baseline; the shaded row is our final model.
Method
StreamVSTI
StreamVSI
VideoLLM-online
6.22
4.94
Dispider
23.39
0.63
Streamo
43.99
21.09
Ours
48.80
29.44
Table 4: Strict-mode average accuracy (%) for streaming methods. The strict protocol scores only the earliest predicted ⟨Response⟩ and counts a missing response as zero, penalizing timing failures directly. Offline models do not respond and are omitted.
Ablations
Δt (s) ↓
Geo. Tokens
RT Loss
All
Future-only
✓
0.18
0.52
✓
✓
0.12
0.35
Table 5: The response time loss targets future questions. Mean response-time error Δt (s, lower is better) over all questions versus current-ask-future questions only, whose answers depend on frames that arrive after the query. Adding the response time loss (row 2) cuts the error most on future questions, where the model must decide how long to wait.
Figure A1: Interface for the human verification study of the visibility annotations. For each of the 500 sampled object–frame pairs, annotators see the projected 3D box overlay, a zoomed crop, and neighboring frames around the pipeline’s first-seen claim, and judge whether the object is genuinely recognizable; the resulting human–pipeline agreement is 84.4% .
Appendix figures & tables4 assets
Supplementary material from the paper’s appendix.
Appendix
Hyperparameter
Value
LoRA rank / α / dropout
128 / 256 / 0.05
LoRA targets
q,k,v,o,gate,up,down
Fully trained modules
embed_tokens , lm_head
Trainable parameters
1,461.2 M ( 15.0% )
Base / projector learning rate
1×10−5 / 5×10−4
Optimizer
AdamW, β=(0.9,0.95)
Appendix
Table A1: Training hyperparameters of SpaTime.
Split
Future
Current
Past
Train ( 132,568 )
41,196
22,268
69,104
Test ( 6,042 ) ∗
2,076
2,075
1,891
Appendix
Table A2: Temporal-channel distribution of StreamVSTI-Bench.
Channel
Bucket
n
Acc. (%)
Exact round
Δ round
Current
MC
1,440
62.50
99.1%
+0.001
Current
Num.
635
35.04
98.7%
+0.005
Future
MC
1,512
59.92
60.4%
−0.371
Future
Num.
564
35.32
74.3%
−0.242
Past
MC
1,346
63.45
100.0%
0
Past
Num.
545
36.90
99.8%
0
Appendix
Table A3: Per-temporal-channel breakdown on StreamVSTI-Bench. “Exact round” is the fraction of samples whose ⟨Response⟩ lands on precisely the ground-truth round; Δ round is the mean of tpred−tgt , so negative values mean the model answers early.
Task / channel
Scene
Question (abridged)
GT (round, ans.)
Pred. (round, ans.)
Rel. pos. (lr) / future
scene0700_02
At 33.40s, will telephone be to the [Left/Right] relative to keyboard? (asked at round 14)
(34, A)
(34, A)
Rel. dist. (v3) / future
scene0580_00
Measuring from the closest point of each object at 43.42s, which of (nightstand, bed) will be closest to the camera? (asked at round 24)
(44, B)
(44, B)
Displacement / future
scene0664_00
How far (in meters) will the camera move between 13.12s and 39.35s? (asked at round 14)
(39, 0.5)
(39, 0.5)
Rel. pos. (ud) / current
scene0231_02
At current time, relative to backpack, is window to the [Up/Down]?
(4, A)
(4, A)
Rel. dist. (v2) / current
scene0580_01
Measuring from the closest point of each object at current time, which of (backpack, table, bed) is closest to the camera?
(39, C)
(39, C)
Movement dir. / past
scene0050_00
Looking back, what was the primary consistent direction of the camera’s movement from 12.11s to 84.80s?
(109, C)
(109, C)
Appendix
Table A4: Qualitative examples across task families and temporal channels on StreamVSTI-Bench. For every example, the predicted ⟨Response⟩ round matches the ground-truth round exactly and the answer is correct; future-channel examples require holding ⟨Standby⟩ for 20 – 25 rounds before responding.
Despite advances in 3D scene understanding, existing 3D Large Multimodal Models operate in offline settings, requiring complete scene observations or predefined video clips. In this paper, we present an online 3D vision-language model that enables real-time spatial understanding from streaming video. Our approach adopts an autoregressive streaming control modeling based on the LLM's next-token prediction objective to learn when to respond, and employs a lightweight Visual-Spatial Feature Integration (VSFI) module to incrementally inject temporally aligned geometry priors into the visual stream. To alleviate long-context decoding overhead, we propose a plug-and-play Geometry-Adaptive Voxel Compression (GAVC) module for efficient visual token compression. To address the scarcity of streaming 3D-language data, we further develop a scalable data generation pipeline that curates over 1M online spatio-temporal 3D QA pairs and establishes a comprehensive benchmark spanning 29 tasks. Extensive experiments show that our approach significantly outperforms both proprietary and open-source models across online and offline 3D spatial understanding, reasoning, and grounding tasks. The project page is available at https://stream3d-vlm.github.io/
Hanxun Yu, Xuan Qu, Lei Ke +4
1Zhejiang University · 2Tencent Hunyuan · 3HKUST +1
We introduce EgoSAT, the first comprehensive benchmark for egocentric video reasoning in streaming settings, designed to evaluate the capabilities of modern vision-language models (VLMs). The benchmark targets streaming interaction understanding, where video frames arrive sequentially and models must continuously interpret evolving visual context. EgoSAT unifies several previously distinct tasks within a single streaming framework. In this formulation, queries about completed events correspond to retrospective reasoning, queries about ongoing activities require online understanding, and queries about future actions involve prospective anticipation. This unified setting requires models to reason about the past, present, and future while operating under the constraint that only previously observed frames are available. EgoSAT contains 1,997 unique videos spanning 165 hours of egocentric footage and around 4,800 high-quality question-answer pairs, carefully designed to probe reasoning across varying temporal contexts. Using this benchmark, we evaluate a diverse set of both open-weight and closed-weight VLMs, providing a systematic assessment of their ability for streaming interaction understanding. By distinguishing answerability and conducting diagnostics on confidence of models, we find existing models not only struggle with prospective and retrospective modeling, but also exhibit severe mis-calibration: confidence often fails to track inherent answerability, leading to dangerous "confidently wrong" behaviors. Project page: https://leiyj23.github.io/EgoSAT/
Yijia Lei, Jinzhao Li, Yichi Zhang +3
College of AI, Tsinghua University · University of Wisconsin–Madison
Vision-Language Models (VLMs) perform well on commonsense reasoning tasks but struggle with visual spatial reasoning. Most existing solutions introduce extra 3D prior inputs or external spatial encoders, which increase complexity and degrade the underlying VLMs' general-purpose capabilities after spatial fine-tuning. To this end, we propose a parameter-efficient \textit{\textbf{Spatio}-vision \textbf{L}anguage \textbf{M}odels (SpatioLM)}, that enhances spatial intelligence without extra 3D prior inputs or third-party spatial encoders. Concretely, we design a plug-and-play and non-invasive spatio-vision module that elicits the spatial knowledge inherent in VLMs. Furthermore, we innovatively leverage pseudo depth and camera information as supervision to guide the model in learning physically coherent representations. Extensive experiments show that SpatioLM achieves significant improvements in diverse tasks, including spatial perception and understanding while effectively limiting the degradation of general capabilities. Notably, the model achieves an impressive score of 71.6 on the VSI-Bench (the first model to surpass 70). In addition, it attains competitive performance when transferred to embodied manipulation tasks. Code is available at \faGithub~spatio-lm.
Jing Wu, Jianhua Wu, Jiayi Guan +5
Xiaomi EV, Beijing, China · College of Automotive and Energy Engineering, Tongji University, Shanghai, China · Independent Researcher