Spatio-temporal video grounding (STVG) aims to localize objects or events described by natural language queries in both space and time. Existing STVG models are typically trained and evaluated under the assumption that each query is relevant to the input video. In this work, we challenge this assumption by studying the behavior of state-of-the-art STVG models under irrelevant queries and missing textual input. Our experiments show that current models can still produce plausible spatio-temporal predictions even when the query is unrelated to the video or removed entirely. We further analyze HCSTVG-v2 and VidSTG to identify dataset regularities that may encourage such query-insensitive behavior. Our study highlights an underexplored limitation of STVG models and motivates negative-aware evaluation protocols and architectures that explicitly assess query relevance.
Figures & tables
Figure 1: Temporal analysis of HCSTVG-v2 and VidSTG: moment duration and its time center. Values are decimal units normalized to the video duration.
Figure 2: Comparison of spatial annotations between HCSTVG-v2 and VidSTG datasets. Datapoints show average bounding box centers on X and Y axes, normalized to video frame size. Scatter plot indicates frequency of centers at the same spot.
Figure 3: Comparison of words distribution in datasets.
HCSTVG-v2
VidSTG declarative
VidSTG interrogative
TubeDETR
TA-STVG
TubeDETR
TA-STVG
TubeDETR
TA-STVG
Query type
m_vIoU
m_tIoU
m_vIoU
m_tIoU
m_vIoU
m_tIoU
m_vIoU
m_tIoU
m_vIoU
m_tIoU
m_vIoU
m_tIoU
Positive
0.343
0.525
0.335
0.548
0.283
0.468
0.342
0.517
0.244
0.461
0.291
0.501
Negative in-domain
0.170
0.416
0.206
0.476
0.083
0.298
0.097
0.314
0.092
0.318
0.101
0.323
-50%
-21%
-39%
-13%
-71%
-36%
-72%
-39%
-62%
-31%
-65%
-36%
Negative out-of-domain
0.161
0.398
0.186
0.457
0.078
0.273
0.106
0.274
0.074
0.276
0.085
0.279
Table 1: Inference results: we evaluate TubeDETR and TA-STVG on VidSTG and HCSTVG-v2 datasets, using m_vIoU and m_tIoU. Three query types are compared: positive, negative in-domain, and negative out-of-domain. For negative queries, we report relative IoU changes compared to positive ones. A smaller decrease in the IoU value indicates a behavior that is less sensitive to changes in the query.
HCSTVG-v2
VidSTG
Query
m_vIoU
m_tIoU
m_vIoU
m_tIoU
Original
0.357±0.01
0.535±0.00
0.267±0.02
0.468±0.01
Empty
0.262±0.01
0.485±0.01
0.151±0.02
0.399±0.00
-27%
-9%
-43%
-15%
Table 2: Training validation results of TA-STVG model on HCSTVG-v2 and VidSTG datasets. Comparison of trainings with original and empty query. For empty query, we report IoU changes relative to original one.
Model version
m_vIoU
m_tIoU
TA-STVG, baseline
0.357±0.01
0.535±0.00
TA-STVG, text-processing modules removed
0.238±0.01
0.479±0.00
Table 3: Results for TA-STVG model trained on HCSTVG-v2 without the text-processing modules.
HCSTVG-v2
VidSTG
m_vIoU
m_tIoU
m_vIoU
m_tIoU
Center-prior baseline
0.063
0.248
0.014
0.089
Table 4: Results from testing evaluation with constant spatio-temporal annotations taken from data distribution charts.
Appendix figures & tables2 assets
Supplementary material from the paper’s appendix.
Appendix
Figure 4: Comparison of averaged spatial heatmaps between HCSTVG-v2 and VidSTG dataset. Heatmaps illustrate mean bounding boxes of each video in the whole dataset.
Figure 5: Qualitative results of spatio-temporal grounding prediction of TA-STVG model (weights shared by its authors) on the example HCSTVG-v2 videos. The green border marks temporal prediction frame, the yellow box stands for spatial prediction. Input queries are in order: (1) positive (2) negative in-domain (3) negative out-of-domain.