Video temporal grounding (VTG) aims to localize the video interval corresponding to a language query. Recent large vision-language models (LVLMs) show great potential in solving such a multi-modal reasoning task. However, long videos often contain large amounts of redundant information that disturbs LVLMs to mine query-relevant evidence. Instead of dense frame sampling which incurs prohibitive training memory, previous reinforcement learning with verifiable rewards (RLVR) works typically utilize sparse sampling, which makes training feasible but may miss critical evidence. In this paper, we propose a summarize before grounding'' framework (named SumGround'') for long-video temporal grounding. The key of SumGround is to perform query-guided chunk condensation to aggregate and retrieve query-relevant evidence. Specifically, we split the video into several chunks and perform two-level chunk condensation. First, we introduce query-guided latent summaries, which is represented as KV states of query-guided prompts, to compress redundant visual tokens into compact query-relevant chunk summaries. Furthermore, we design an associative summary retrieval scheme to rank and select chunk summaries that are most likely to contain the event interval. Both query-guided latent summary and associative summary retrieval schemes are enabled by RLVR. To reduce memory consumption, we propose a length-aware gradient gating module to selectively stop gradient back-propagated to visual tokens. Extensive experiments demonstrate that SumGround performs favorably against previous state-of-the-art methods across multiple downstream datasets, with remarkable gains on long videos.
Figures & tables
Figure 1: Overview of the motivation and SumGround. (a) Dense sampling retains fine-grained evidence but also introduces extensive query-irrelevant content and prohibitive training memory, whereas sparse sampling may miss key evidence. (b) SumGround processes the video chunk by chunk, compresses query-relevant evidence into sequential summary states, and retains query-associated summaries for final grounding. We gate visual-token gradients according to video length while continuing to optimize summary accumulation, retrieval, and grounding.
Figure 2: Framework of SumGround. (a) For each video chunk Ci , the LVLM jointly writes a query-guided summary state Si and estimates its query-association score Δi , conditioned on previously retained summaries. Only the system and summary KV states are propagated to subsequent chunks, while dense visual states and the states used for association estimation are discarded. (b) After all chunks are processed, SumGround selects an anchor and retrieves a contiguous set of query-associated summaries according to their association scores. The retrieved summaries are used to generate the final temporal grounding boundaries.
Method
Long-video Benchmarks
Standard-video Benchmarks
Ego4D-NLQ
TACoS
Charades-STA
ActivityNet-Captions
R1@.3
R1@.5
mIoU
R1@.3
R1@.5
mIoU
R1@.5
R1@.7
mIoU
R1@.5
R1@.7
mIoU
Supervised / SFT-based grounding methods
UniVTG ( Lin et al. 2023 )
6.48
3.48
4.63
5.17
1.27
4.40
25.22
10.03
27.12
11.10
4.06
16.86
VTG-LLM ( Guo et al. 2025 )
1.71
0.46
1.36
6.87
2.92
5.27
34.11
15.81
34.93
12.32
6.74
17.86
TimeSuite ( Zeng et al. 2024 )
0.88
0.43
0.94
6.75
2.50
5.71
48.95
24.65
45.91
16.56
9.28
22.03
Table 1: Comparison with previous methods on four held-out benchmarks. Ego4D-NLQ and TACoS are long-video benchmarks with average durations of 500s and 368s, respectively, while Charades-STA and ActivityNet-Captions with average durations of 29s and 118s. † denotes methods retrained with the same training data and evaluation pipeline as SumGround. Gray-shaded entries denote results reproduced or evaluated by us. We use officially released checkpoints and codes. Other numbers are directly cited from previous papers. SumGround w/o. Retrieval means we remove the associative summary retrieval operation from our method. Best and second-best results are shown in bold and underlined , respectively.
Setting
R1@0.3
R1@0.5
R1@0.7
mIoU
(a) Summary Type
w/o. Summary
78.09
63.01
36.42
54.46
<summary>
73.12
54.38
27.28
48.78
Query-Agnostic
76.96
60.56
34.81
52.93
Query-Guided (Ours)
81.83
69.57
44.30
58.75
(b) Chunk-wise Dependency
Table 2: Ablation of summary design on Charades-STA. All variants use approximately the same retained token budget. (a) Summary representation comparison The “w/o. Summary” directly grounds from the visual inputs without any summary operations. “ <summary> ” means learned <summary> tokens. “Query-Agnostic” means we use a general summary prompt which does not rely on the query content. To simplify our comparison, we treat the whole video as a single chunk. (b) Chunk-wise Dependency comparison. “Independent” means the summary of each chunk is independent on those of other chunks.
Retrieval Strategy
R1@0.3
R1@0.5
R1@0.7
mIoU
Random
28.32
19.57
9.75
19.46
Anchor Only
43.64
30.17
15.20
29.67
Ours
52.96
38.07
18.67
35.93
Table 3: Comparison of summary retrieval strategies on TACoS. Random retains the same number of summaries as our method but selects them at random. Anchor Only retains only the anchor summary.
γ
Long Video
R1@0.3
R1@0.5
mIoU
1
✓
OOM
OOM
OOM
0
✓
27.44
14.32
20.42
1
✗
45.66
29.94
31.46
I[N≤W]
✓
52.96
38.07
35.93
Table 4: Comparison of gradient-gating configurations on TACoS. γ controls whether gradients propagate through visual-token states ( γ =1 and 0 means retaining and stopping visual gradients respectively), and Long Video means long videos with N>W are utilized in training. The total number of training samples is kept the same across all variants. The last row denotes length-aware gradient gating. OOM denotes out of memory.
Figure 3: Peak GPU memory during training as the video length increases from 15 to 60 frames. All inputs are divided into 15-frame chunks. Full Visual Gradients retains gradients through all visual tokens, while ours applies length-aware gradient gating technique.
Split
Dataset
Video Len.
Moment Len.
Views
Domain
Training
NaQ ( Ramakrishnan et al. 2023 )
413s
1.1s
Ego
Open
DiDeMo ( Anne Hendricks et al. 2017 )
29s
7.5s
Exo
Open
QuerYD ( Oncescu et al. 2021 )
278s
13.6s
Ego & Exo
Open
HiRest ( Zala et al. 2023 )
263s
18.9s
Ego & Exo
Open
COIN ( Tang et al. 2019 )
145s
14.9s
Ego & Exo
Open
Momentor ( Qian et al. 2024 )
403s
49.5s
Ego & Exo
Open
Table 5: Statistics of temporal grounding datasets used in our experiments. “Video Len.” and “Moment Len.” denote the average video duration and average annotated temporal moment duration, respectively.
Method
Tvis
Δ Mem. (GB) ↓
Latency (s) ↓
R@0.3 ↑
R@0.5 ↑
mIoU ↑
Uniform-Single
3,829
1.50
1.17
2.49
1.33
1.77
Full-Dense
51,463
19.75
17.58
2.18
1.02
1.74
SumGround
51,463
3.95
14.78
16.42
9.51
11.83
Table 6: Efficiency–accuracy trade-off on Ego4D-NLQ. Tvis denotes the average cumulative number of visual tokens processed per sample; for SumGround, it is summed over all chunk-wise passes. Δ Mem. denotes the increase in peak GPU memory relative to the approximately 15.5 GB model-only footprint measured under the same inference setup.
Dataset
Method
mIoU
R1@0.3
R1@0.5
R1@0.7
Charades-TimeLens
Time-R1
36.6
57.9
32.0
16.9
TimeLens
48.8
70.5
55.6
28.4
SumGround
41.94
62.09
40.11
20.43
ActivityNet-TimeLens
Time-R1
33.1
44.8
31.0
19.0
TimeLens
46.2
62.8
51.0
32.6
SumGround
43.92
59.71
45.88
29.71
Table 7: Additional evaluation on the TimeLens benchmark. Best results are shown in bold , and second-best results are underlined .
Method
Ego4D-NLQ
TACoS
GT-Contain
GT-Overlap
GT-Contain
GT-Overlap
Qwen2.5-VL
0.43
0.46
0.67
0.77
UniTime
0.43
0.50
0.72
0.89
SumGround
0.52
0.55
0.78
0.88
Table 8: Explicit retrieved-context coverage on long-video temporal grounding benchmarks. GT-Contain measures whether the retrieved source region fully contains the annotated interval, while GT-Overlap measures whether the two have any temporal intersection. Values are proportions, and higher is better.
Dataset
Variant
R@0.3
R@0.5
R@0.7
mIoU
TACoS
w/o Auxiliary Cache
47.14
29.59
12.67
31.45
SumGround
52.96
38.07
18.67
35.93
Charades-STA
w/o Auxiliary Cache
77.07
61.45
34.17
52.87
SumGround
79.54
65.91
39.57
56.10
Table 9: Ablation of the auxiliary local visual cache on TACoS and Charades-STA. The ablated variant sets C=∅ throughout training and inference, while all other settings remain unchanged.
Figure 4: Illustration of prompts at both training and inference time.
Figure 5: Case study of anchor chunk selection on Ego4D-NLQ. Orange denotes the selected anchor chunk, while green denotes the annotated ground-truth span. Case (c) is a genuine selection error, where the selected chunk contains a similar action but the wrong object. In contrast, Cases (a) and (b) show semantically relevant anchor chunks that match the query but fall outside the annotated ground-truth span, suggesting that some measured failures may reflect to repeated or under-annotated query-relevant events.
Current Video-LLM approaches for Video Temporal Grounding (VTG) typically rely on direct timestamp generation from an unstructured visual-token stream, often leading to brittle numerics and inconsistent boundaries. To address this, we propose Foresee-to-Ground (F2G), a framework that reformulates VTG as a verifiable Identify-then-Measure problem. F2G integrates Predictive Temporal Perception with Evidence-Driven Reasoning: it learns boundary-sensitive temporal representations to build a video-wide evidence pool of candidate event segments, and exposes these segments to the LLM as citable evidence units that bind boundary prediction to explicit event hypotheses. By decoupling event identification from precise boundary measurement, F2G stabilizes grounding and makes predictions verifiable. Extensive experiments demonstrate that F2G consistently improves grounding accuracy across diverse benchmarks, transfers robustly across different Video-LLM backbones, and preserves general video understanding capabilities.
Zelin Zheng, Xinyan Liu, Ruixin Li +4
School of Computer Science and Technology, University of Chinese Academy of Sciences, Beijing, China · Beijing Key Laboratory of Embodied Intelligence Computing, Beijing, China · Faculty of Computing, Harbin Institute of Technology, Weihai, China +1
Spatio-temporal grounding in long videos requires precise temporal localization and robust object tracking conditioned on natural-language queries. While recent vision-language models (VLMs) show strong reasoning ability, directly applying frame-by-frame inference to long sequences is computationally expensive and unstable. We propose a practical pipeline that shifts from frame-level to second-level tracking and performs cross-second smoothing to preserve continuity while reducing sequence length. To improve reasoning supervision, we synthesize chain-of-thought style trajectories using advanced multimodal models for temporal localization and target selection, and replace generated spatio-temporal coordinates with ground-truth annotations to avoid noisy supervision. We further optimize the policy with reinforcement learning using a verifier based on t_IoU+mv_IoU. Experiments across multiple FPS settings show that our method achieves a strong trade-off between efficiency and localization quality.
Video Temporal Grounding (VTG) localizes the temporal boundaries of query-relevant moments in long, untrimmed videos, making video-language-model prohibitively expensive. While recent training-free token pruning has shown success in video question answering, naively applying these objectives to VTG causes drastic degradation, as VTG crucially depends on boundary-sensitive evidence and cross-frame reasoning chains. We therefore identify two VTG-specific pruning principles: evidence retention, which keeps query-critical patches especially around event boundaries, and connectivity strength, which preserves cross-frame connectivity for long-range evidence aggregation. Building on these insights, we propose SemVID, a training-free pruning framework that constructs a compact yet coherent token subset with complementary semantic roles. SemVID first allocates per-frame budgets by balancing query relevance and inter-frame variation to avoid over-pruned segments, and then selects three types of tokens: object tokens for diverse query-critical evidence, motion tokens to capture meaningful transitions and serve as cross-frame relays, and context tokens for scene continuity. Extensive experiments show that SemVID achieves a strong accuracy-efficiency trade-off, retaining up to 95.4% mIoU with only 12.5% visual tokens and delivering up to a 5.8x prefill speedup, consistently outperforming prior methods under the same budgets. Our code is available at https://github.com/JiaqiLi404/SemVID
Jiaqi Li, Shuntian Zheng, Yixian Shen +4
University of Warwick, Coventry CV4 7AL, United Kingdom · University of Amsterdam, Amsterdam 1012 WX, Netherlands · Amazon AGI, The United States of America