cs.CVOct 5, 2026

Investigating Query-Insensitive Behavior in Spatio-Temporal Video Grounding

Authors: Eryk Kołodziejczyk, Alberto Presta, Karol Szurkowski, Michal Byra

Organizations: Samsung AI Center, Warsaw, Poland · Institute of Fundamental Technological Research, Polish Academy of Sciences, Warsaw, Poland

Abstract

Spatio-temporal video grounding (STVG) aims to localize objects or events described by natural language queries in both space and time. Existing STVG models are typically trained and evaluated under the assumption that each query is relevant to the input video. In this work, we challenge this assumption by studying the behavior of state-of-the-art STVG models under irrelevant queries and missing textual input. Our experiments show that current models can still produce plausible spatio-temporal predictions even when the query is unrelated to the video or removed entirely. We further analyze HCSTVG-v2 and VidSTG to identify dataset regularities that may encourage such query-insensitive behavior. Our study highlights an underexplored limitation of STVG models and motivates negative-aware evaluation protocols and architectures that explicitly assess query relevance.

Figures & tables

Appendix figures & tables2 assets

Supplementary material from the paper’s appendix.

Appendix

Explore similar work

CardsList
  1. Learning Compositional Spatio-Temporal Video Grounding with Synthetic Curriculum

    Aug 31, 2026Xingjian Wang, Shijian Wang, Yibo Wang +4Video Temporal GroundingSynthetic Training Data

  2. DSE-VTG: Dual-Side Enhancement for Training-Free Video Temporal Grounding

    Sep 8, 2026Zhuo Cao, Bingqing Zhang, Sen Wang +1Video Temporal GroundingDense Temporal Annotation

  3. DART: Difficulty-Adaptive Routing for Zero-Shot Video Temporal Grounding

    Jul 1, 2026Zhengbo Zhang, Mark He Huang, Zhigang Tu +1Video Temporal GroundingTemporal Grounding