cs.CVOct 6, 2026

SpaTime: Streaming Vision-Language Models for Spatio-temporal Reasoning

Authors: Hairong Yin, Huangying Zhan, Shin-Fang Chng, Yi Xu, Raymond A. Yeh

Organizations: Department of Computer Science, Purdue University, USA · Goertek Alpha Labs, USA

Abstract

Embodied agents must reason about 3D space while the video is still arriving, answering questions as soon as they have observed enough of the scene. VLMs that incorporate 3D geometric priors achieve strong spatial reasoning, but they operate offline, i.e., the full video must be available before they produce an answer. Streaming VLMs process frames causally and decide for themselves when to respond, yet they lack explicit 3D representations. We present SpaTime, a streaming VLM that fuses causal geometry tokens into the language model at every frame, using only the frames observed so far. To supervise when the model answers, we propose a response-time loss that maps per-frame response probabilities to a differentiable expected response time and penalizes the distance from the ground-truth frame. For evaluation, we construct StreamVSTI-Bench and StreamVSI-Bench, streaming adaptations of VSTI-Bench and VSI-Bench. On StreamVSTI-Bench, SpaTime reaches 49.2% overall accuracy and reduces the mean response-time error by 66% relative to the strongest streaming baseline.

Figures & tables

Appendix figures & tables4 assets

Supplementary material from the paper’s appendix.

Appendix

Explore similar work

CardsList
  1. Stream3D-VLM: Online 3D Spatial Understanding with Incremental Geometry Priors

    Jun 5, 2026Hanxun Yu, Xuan Qu, Lei Ke +43D Spatial ReasoningVoxel

  2. EgoSAT: A Comprehensive Benchmark of Egocentric Streaming Interaction Understanding

    Jun 23, 2026Yijia Lei, Jinzhao Li, Yichi Zhang +3Egocentric VideoQuestion-Answer Pairs

  3. SpatioLM: Towards General Physical Spatial Intelligence in Vision-Language Models

    Aug 3, 2026Jing Wu, Jianhua Wu, Jiayi Guan +5Recent Vision-Language ModelsStable Spatial Understanding