Precise temporal grounding requires distinguishing when a queried event occurs from when its participating entities are merely visible. We propose Diffusion-Grounded VideoLLM, which conditions temporal feature extraction on query-relevant entities before language reasoning. The framework tracks entities named in the query and uses their masks to condition a frozen video diffusion backbone. Intermediate spatiotemporal features are extracted through truncated denoising and combined with entity tokens and timestamp embeddings. The language model uses this evidence together with the full query to generate temporal intervals and answers to grounded questions. On Charades-STA and NExT-GQA, the model obtains 43.5 mIoU and 28.4 Acc@GQA, improving the reported Grounded-VideoLLM reference by 6.7 and 1.7 points, respectively. Component and entity-pathway ablations support the usefulness of conditioning diffusion features on query-relevant entities for temporal grounding.
Figures & tables
Figure 1: Overview of Diffusion-Grounded VideoLLM. Query-derived tracks supply entity tokens and masks that condition VACE. Early diffusion features receive alignment supervision during training. Visual, text, and diffusion features are concatenated with timestamp information for interval and answer generation. Token order is schematic.
Model
Size
Charades-STA
DiDeMo
R.3
R.5
R.7
mIoU
R.3
R.5
R.7
mIoU
Video-LLaMA [ 25 ]
7B
25.2
10.6
3.4
16.8
20.1
8.2
2.5
14.3
SeViLA [ 22 ]
3B
27.0
10.5
5.8
18.3
23.5
9.8
3.6
15.9
Video-ChatGPT [ 12 ]
7B
27.2
6.2
1.9
19.7
19.8
6.5
1.2
13.7
Momentor [ 13 ]
7B
42.6
26.6
11.6
28.5
38.2
21.8
9.4
26.5
VTimeLLM [ 5 ]
7B
51.0
27.5
11.4
31.2
45.0
28.8
12.0
27.9
Table 1: Temporal grounding on Charades-STA and DiDeMo. R@1 is reported at different IoU thresholds.
Model
Acc@GQA
mIoP
IoP@0.5
mIoU
IoU@0.5
VIOLETv2 [ 1 ]
12.8
23.6
23.3
3.1
1.3
Temp[CLIP] NG+ [ 21 ]
16.0
25.7
25.5
12.1
8.9
SeViLA [ 22 ]
16.6
29.5
22.9
21.7
13.8
LangRepo [ 7 ]
17.1
31.3
28.7
18.5
12.2
FrozenBiLM NG+ [ 21 ]
17.5
24.2
23.7
9.6
6.1
VideoStreaming [ 14 ]
17.8
32.2
31.0
19.3
13.3
Table 2: Grounded VideoQA results on NExT-GQA.
Model
MSVD-QA
MSRVTT-QA
ANet-QA
Acc.
Score
Acc.
Score
Acc.
Score
Video-LLaMA [ 25 ]
51.6
2.5
29.6
1.8
12.4
1.1
Video-ChatGPT [ 12 ]
64.9
3.3
49.3
2.8
35.2
2.7
VideoChat2 [ 9 ]
70.0
3.9
54.1
3.3
49.1
3.3
ST-LLM [ 10 ]
74.6
3.9
63.2
3.4
50.9
3.3
Momentor [ 13 ]
68.9
3.6
55.6
3.0
40.8
3.2
Table 3: Open-ended VideoQA results. We report answer accuracy (Acc.) and GPT-based evaluation score.
Variant
Diff.
Entity
Mix.
R@0.3
R@0.5
R@0.7
mIoU
Base
54.2
36.4
19.7
36.8
+ Diffusion
✓
56.1
41.3
20.9
39.5
+ Entity Conditioning
✓
✓
57.3
43.2
22.1
42.1
+ Mixed Tokens
✓
✓
✓
58.7
45.2
23.0
43.5
Table 4: Cumulative component ablation on Charades-STA.
Object Tokens
Mask Conditioning
R@0.3
R@0.5
R@0.7
mIoU
56.1
41.3
20.9
39.5
✓
56.6
41.9
21.3
40.3
✓
57.0
42.6
21.8
41.5
✓
✓
57.3
43.2
22.1
42.1
Table 5: Entity-conditioning pathway ablation on Charades-STA.
Extraction Point
R@0.3
R@0.5
R@0.7
mIoU
10%
58.7
45.2
23.0
43.5
30%
57.2
41.9
22.0
41.4
50%
56.7
40.6
20.8
40.3
Table 6: Effect of diffusion feature extraction point on Charades-STA.
Tracking Strategy
R@0.3
R@0.5
R@0.7
mIoU
Frame-wise grounding
56.1
40.0
21.1
39.9
Tracking, K=1
57.0
41.8
21.8
41.2
Tracking, K=2
58.0
43.6
22.5
42.5
Tracking, K=3
58.7
45.2
23.0
43.5
Table 7: Effect of persistent entity tracking on Charades-STA.
Setting
Value
R@0.3
R@0.5
R@0.7
mIoU
Nt
4
57.0
41.6
21.6
41.2
8
58.7
45.2
23.0
43.5
16
58.1
44.1
22.6
42.8
No
2
57.5
42.9
22.0
42.0
4
58.7
45.2
23.0
43.5
8
58.3
44.5
22.7
43.0
Table 8: Effect of representation budget on Charades-STA.
Video Temporal Grounding (VTG) localizes the video segment that matches a natural-language query. Many queries describe an action performed by a particular entity. Existing methods often encode the query as a whole or use general video-text interactions, without explicitly checking whether the action and entity occur together. They may therefore select a segment that contains both concepts but not the event described by the query. We propose Interaction Aligned Action-Entity Video Temporal Grounding (IAE-VTG), which models this rela?tionship at both the representation and training assignment levels. First, the Fine-grained Disentangled Interaction Module (FDIM) separates action and entity related query information and aligns it with complementary motion and appearance features. It then combines token-level interactions to build representations that capture the relationship between the action and entity. Second, Interaction-Sensitive Assignment (ISA) adds this interaction evidence to bipartite matching, so training targets are selected using both temporal overlap and semantic compatibility. This reduces supervision from temporally plausible but semantically incorrect proposals. Experiments on QVHighlights, Charades?STA, and TACoS show that IAE-VTG consistently improves strong baselines and achieves competitive or state-of-the-art performance on standard grounding metrics. Additional analyses show that the method is especially effective when similar actions or entities appear at multiple times and produces more reliable assignments for complex events.
Shiwen Zhao, Qi Zhang, Sezer Karaoglu +2
School of Computer Science, The University of Sydney, Sydney, NSW 2006, Australia · Informatics Institute at University of Amsterdam, 1098 XH Amsterdam, the Netherlands
Current Video-LLM approaches for Video Temporal Grounding (VTG) typically rely on direct timestamp generation from an unstructured visual-token stream, often leading to brittle numerics and inconsistent boundaries. To address this, we propose Foresee-to-Ground (F2G), a framework that reformulates VTG as a verifiable Identify-then-Measure problem. F2G integrates Predictive Temporal Perception with Evidence-Driven Reasoning: it learns boundary-sensitive temporal representations to build a video-wide evidence pool of candidate event segments, and exposes these segments to the LLM as citable evidence units that bind boundary prediction to explicit event hypotheses. By decoupling event identification from precise boundary measurement, F2G stabilizes grounding and makes predictions verifiable. Extensive experiments demonstrate that F2G consistently improves grounding accuracy across diverse benchmarks, transfers robustly across different Video-LLM backbones, and preserves general video understanding capabilities.
Zelin Zheng, Xinyan Liu, Ruixin Li +4
School of Computer Science and Technology, University of Chinese Academy of Sciences, Beijing, China · Beijing Key Laboratory of Embodied Intelligence Computing, Beijing, China · Faculty of Computing, Harbin Institute of Technology, Weihai, China +1
Temporal grounding--returning the interval [ts,te] for a natural-language query over a video--is the language interface to long-form video, yet has been studied on short videos; the dynamics of hour-scale natural-language grounding remain underexplored. We take the position that at hour-scale, the binding constraint is search, not recognition: Video-LLMs are bottlenecked not by localizing a nearby event, but--given a natural-language query--by searching for the relevant region of a long video. To test this, we release ExtremeWhenBench, the first open hour-scale grounding benchmark (2,273 queries over 194 videos, mean 75.7 min, max 9 hr) with an open-form query distribution. Every open Video-LLM collapses while a frame-level retrieval baseline outperforms them; a failure taxonomy attributes 85% of failures to search; and a retrieve-then-ground hybrid recovers 6.7x over the monolithic Video-LLM--mirroring retrieve-then-read in open-domain QA.