Spatio-temporal video grounding (STVG) aims to localize objects or events described by natural language queries in both space and time. Existing STVG models are typically trained and evaluated under the assumption that each query is relevant to the input video. In this work, we challenge this assumption by studying the behavior of state-of-the-art STVG models under irrelevant queries and missing textual input. Our experiments show that current models can still produce plausible spatio-temporal predictions even when the query is unrelated to the video or removed entirely. We further analyze HCSTVG-v2 and VidSTG to identify dataset regularities that may encourage such query-insensitive behavior. Our study highlights an underexplored limitation of STVG models and motivates negative-aware evaluation protocols and architectures that explicitly assess query relevance.
Figures & tables
Figure 1: Temporal analysis of HCSTVG-v2 and VidSTG: moment duration and its time center. Values are decimal units normalized to the video duration.
Figure 2: Comparison of spatial annotations between HCSTVG-v2 and VidSTG datasets. Datapoints show average bounding box centers on X and Y axes, normalized to video frame size. Scatter plot indicates frequency of centers at the same spot.
Figure 3: Comparison of words distribution in datasets.
HCSTVG-v2
VidSTG declarative
VidSTG interrogative
TubeDETR
TA-STVG
TubeDETR
TA-STVG
TubeDETR
TA-STVG
Query type
m_vIoU
m_tIoU
m_vIoU
m_tIoU
m_vIoU
m_tIoU
m_vIoU
m_tIoU
m_vIoU
m_tIoU
m_vIoU
m_tIoU
Positive
0.343
0.525
0.335
0.548
0.283
0.468
0.342
0.517
0.244
0.461
0.291
0.501
Negative in-domain
0.170
0.416
0.206
0.476
0.083
0.298
0.097
0.314
0.092
0.318
0.101
0.323
-50%
-21%
-39%
-13%
-71%
-36%
-72%
-39%
-62%
-31%
-65%
-36%
Negative out-of-domain
0.161
0.398
0.186
0.457
0.078
0.273
0.106
0.274
0.074
0.276
0.085
0.279
Table 1: Inference results: we evaluate TubeDETR and TA-STVG on VidSTG and HCSTVG-v2 datasets, using m_vIoU and m_tIoU. Three query types are compared: positive, negative in-domain, and negative out-of-domain. For negative queries, we report relative IoU changes compared to positive ones. A smaller decrease in the IoU value indicates a behavior that is less sensitive to changes in the query.
HCSTVG-v2
VidSTG
Query
m_vIoU
m_tIoU
m_vIoU
m_tIoU
Original
0.357±0.01
0.535±0.00
0.267±0.02
0.468±0.01
Empty
0.262±0.01
0.485±0.01
0.151±0.02
0.399±0.00
-27%
-9%
-43%
-15%
Table 2: Training validation results of TA-STVG model on HCSTVG-v2 and VidSTG datasets. Comparison of trainings with original and empty query. For empty query, we report IoU changes relative to original one.
Model version
m_vIoU
m_tIoU
TA-STVG, baseline
0.357±0.01
0.535±0.00
TA-STVG, text-processing modules removed
0.238±0.01
0.479±0.00
Table 3: Results for TA-STVG model trained on HCSTVG-v2 without the text-processing modules.
HCSTVG-v2
VidSTG
m_vIoU
m_tIoU
m_vIoU
m_tIoU
Center-prior baseline
0.063
0.248
0.014
0.089
Table 4: Results from testing evaluation with constant spatio-temporal annotations taken from data distribution charts.
Appendix figures & tables2 assets
Supplementary material from the paper’s appendix.
Appendix
Figure 4: Comparison of averaged spatial heatmaps between HCSTVG-v2 and VidSTG dataset. Heatmaps illustrate mean bounding boxes of each video in the whole dataset.
Figure 5: Qualitative results of spatio-temporal grounding prediction of TA-STVG model (weights shared by its authors) on the example HCSTVG-v2 videos. The green border marks temporal prediction frame, the yellow box stands for spatial prediction. Input queries are in order: (1) positive (2) negative in-domain (3) negative out-of-domain.
Despite the impressive progress of recent MLLMs on spatio-temporal video grounding (STVG), existing evaluations and training data focus primarily on simple queries. They largely overlook the compositional queries prevalent in real-world scenarios, where a target must be disambiguated by jointly reasoning about its attributes and relations to other entities. To bridge this gap, we propose Compositional Spatio-Temporal Video Grounding (CompSTVG), a task that requires models to process complex textual queries where every intertwined attribute and relational cue is essential for disambiguation. To facilitate this task at scale, we build a synthetic data engine that leverages a spatio-temporal scene graph as a difficulty measure and casts difficulty-controlled query synthesis as a constraint programming problem, producing difficulty-graded data for both evaluation and training. Built on this engine, we introduce STVG-CompBench, a benchmark stratified by explicit difficulty levels that jointly capture temporal complexity and spatial interference. Evaluating 11 representative STVG models on STVG-CompBench reveals that current models perform poorly on compositional queries, exhibiting a sharp performance drop that is typically obscured by overall dataset-level averages. We further construct synthetic training data and propose CurrSTVG, a curriculum reinforcement learning framework that delivers consistent gains, with the largest improvements observed on the most challenging compositional queries.
Xingjian Wang, Shijian Wang, Yibo Wang +4
1Monash University · 2Southeast University · 3Shanghai University of Electric Power
Text-guided Video Temporal Grounding (VTG) aims to localize the relevant segments in an untrimmed video based on text queries, yet collecting dense temporal annotations and training task-specific models remain costly and brittle under distribution shift. Recent training-free VTG approaches mitigate this issue by directly matching pretrained vision-language representations, but they still face two fundamental information bottlenecks: frame-wise visual encoding overlooks temporal dynamics, while fixed query embeddings cannot resolve query ambiguity. To address these issues, we propose DSE-VTG, a \underline{D}ual-\underline{S}ide \underline{E}nhancement framework that addresses both without any task-specific training. On the visual side, Multi-scale Similarity Fusion (MSF) combines frame- and clip-level similarities into a unified, temporally aware similarity profile. On the textual side, Query-level Test-Time Adaptation (Q-TTA) optimizes a lightweight additive offset to adapt the query embedding to the video at test time, without finetuning the backbone or calling external large language models. Extensive experiments on three standard and two OOD benchmarks show that DSE-VTG achieves state-of-the-art performance among training-free methods. On Charades-STA, it improves mIoU over the strongest prior training-free method by 5.61 points. Under distribution shift, DSE-VTG reaches 50.86 mIoU on Charades-CG Novel-Word, surpassing the strongest supervised baseline by 2.76 mIoU. Our code will be released upon acceptance.
Zero-shot video temporal grounding (VTG) localizes events in untrimmed videos from natural language queries without task-specific training. Existing methods rely on frame-query feature matching, which suffices for simple events but struggles with complex multi-stage queries that require understanding temporal ordering and causal structure -- a disparity we call the reasoning gap. We propose DART (Difficulty-Adaptive Routing for Temporal Grounding), which bridges this gap by coupling difficulty-aware routing with structured reasoning in large vision-language models. A query-conditioned Determinantal Point Process (DPP) serves a dual role: selecting diverse, query-relevant keyframes as temporal evidence, and providing spectral entropy as a difficulty indicator. Simple queries are routed to a Fast path for direct prediction, while complex queries follow a Slow path with Temporal Markup Prompting, which decomposes localization into global event analysis, per-frame temporal role annotation, and boundary extraction. On Charades-STA and ActivityNet Captions, DART achieves state-of-the-art zero-shot performance across both identically distributed and multiple out-of-distribution settings, improving mIoU by up to 3.5 points over the strongest baseline while using over 7 times fewer frames. The project homepage is available at https://dart-vtg.github.io/.
Zhengbo Zhang, Mark He Huang, Zhigang Tu +1
Singapore University of Technology and Design · Wuhan University · University of California, Merced