Video temporal grounding supports applications such as video search, content review, and automated editing by localizing events described in natural language. Yet existing generative models typically output timestamps without explicit interval-level confidence scores to guide candidate selection. We separate candidate generation from acceptance by scoring individual intervals within the original decoding pass. A lightweight confidence head reads pooled decoder states, providing an explicit score trained for interval selection. Offline verifier scores supervise the head on fixed candidate sequences, and temporal-overlap labels adapt it to current rollouts during reinforcement learning. GT-anchored candidate-pool supervision and set-level optimization train the generator. The resulting scores support ranking, threshold-based selection, and rejection without invoking an external verifier at inference. On a fixed OMTG-Bench candidate pool, confidence raises query-macro [email protected] from 9.95% to 14.42% over generation order at a 10% global return budget, and from 26.48% to 31.12% at a 25% budget. The continuous scores let downstream applications adjust return budgets or acceptance thresholds to match their precision-recall preferences, without regenerating candidate intervals.
Figures & tables
Figure 1: Localization alone does not meet downstream selection needs. Applications differ in their tolerance for false and missed matches. Timestamp-only generative TVG lacks explicit control for adapting interval selection to these requirements. Frames are synthetic and scenarios illustrative.
Figure 2: Single-pass grounding with confidence. A span-pooled head scores temporal intervals from the same decoding pass. Generation-order NMS forms pool P ; confidence thresholding selects Sτ . GT-anchored SFT and set-level RL train the generator; offline supervision and online overlap labels train the head. Dashed arrows update decoder θ or head ψ ; numerical examples are illustrative.
Budget
Order
Logp
Verifier
Confidence
10%
9.95
9.41
11.00
14.42
25%
26.48
22.37
26.31
31.12
50%
51.34
43.05
47.72
52.82
75%
64.76
61.00
62.45
65.30
Table 1: Confidence-score comparison at equal global return budgets. Official OMTG-Bench [email protected] (%, ↑ ), averaged over all 320 queries using the same 1,314 candidates. Budget denotes the retained fraction of candidates returned across queries.
Figure 3: Confidence-controlled selection and localization. (a) OMTG-Bench precision–recall curves sweep a global threshold on each model’s fixed candidates after generation-order NMS (0.3), without re-decoding. The high-precision/high-recall region is shown. The gray diamond is GT-only SFT at the official visual/512-token budget, without NMS or confidence selection. (b) Localization F1 versus false-alarm rate on the synthetic matched/unmatched splits of Table 4 . Colors identify datasets; arrows connect native TimeLens2-4B to our RL model, with native TimeLens2-8B also shown. Higher F1 and lower false-alarm rate are better. In both panels, stars mark the global test-F1-optimal operating points reported in the corresponding tables.
Table 2: Multi-occurrence grounding on OMTG-Bench. Official metrics (%). Native TimeLens2 uses the official visual/512-token budget; published baselines retain source protocols ( Zhu et al., 2026 ; Xu et al., 2026 ) . † : one global [email protected] threshold.
Raw
NMS
NMS + confidence †
Model
Recall
tF1
Recall
tF1
Recall
tF1
SFT
67.85
34.65
65.06
64.17
64.92
64.54
RL
73.64
41.39
70.51
67.07
70.10
67.39
Table 3: Candidate coverage and successive selection stages. Official [email protected] and [email protected] (%) on fixed decoded pools. NMS uses generation order at tIoU 0.3; † denotes each model’s global test-F1-optimal confidence threshold.
Table 4: Localization and rejection on balanced matched/synthetic-mismatched splits. All prompts allow empty outputs. † denotes one benchmark-wide test-F1-optimal threshold; native baselines retain their original output rules. [email protected] and false-alarm rate (FAR) are percentages.
Appendix figures & tables10 assets
Supplementary material from the paper’s appendix.
Appendix
Budget
Δ vs. Order
Δ vs. Logp
Δ vs. Verifier
10%
+4.47 [2.30, 6.04]
+5.01 [3.26, 6.36]
+3.42 [1.36, 4.56]
25%
+4.64 [2.26, 7.09]
+8.75 [5.86, 11.29]
+4.81 [2.33, 7.16]
Appendix
Table 5: Official recall differences at tight global budgets. Confidence minus each baseline in query-macro [email protected] (percentage points), with paired 95% percentile intervals.
mIoU
R@1, 0.5
Dataset
Queries
First
Confidence
Δ mIoU [95% CI]
First → Conf.
Charades
3,363
47.89
49.91
+2.02 [1.39, 2.68]
55.16 → 57.21
ActivityNet
4,500
45.96
46.40
+0.44 [ − 0.21, 1.11]
50.76 → 51.33
QVHighlights
1,541
57.26
57.95
+0.68 [ − 0.38, 1.69]
61.19 → 62.10
Appendix
Table 6: Single-answer ranking on fixed decoded pools. Complete matched subsets of the mixed-query TimeLens runs; mIoU and R@1 at tIoU 0.5 are percentages. The paired mIoU changes and 95% intervals use 2,000 source-video bootstrap samples. Both selectors reuse the saved responses to the prompt allowing [] ; invalid or empty answers have zero IoU.
Table 7: Confidence selection after geometric suppression. Official OMTG-Bench metrics (%), averaged over all 320 queries. Confidence uses each model’s global F1-optimal test-oracle threshold; NMS alone retains every survivor.
Variant
Recall @0.5
C-Acc
tF1 @0.5
tIoU
EtF1
Candidate generation: raw outputs
GT-only SFT
61.55
48.13
62.59
58.67
38.08
Ours (SFT), raw
67.85
0.00
34.65
53.35
0.00
Ours (RL), raw
73.64
0.31
41.39
57.21
0.01
Final selection: NMS + confidence selection
Ours (SFT) †
64.92
51.88
64.54
60.75
40.26
Appendix
Table 8: Candidate generation and final selection on OMTG-Bench. All rows use the official visual budget and greedy decoding with 512 output tokens on all 320 queries (1,173 GT intervals); metrics are percentages. The upper block shows raw outputs; the lower applies generation-order NMS (0.3) and confidence selection. † denotes one global F1-optimal test-oracle threshold, shared by all metrics in that row.
Table 9: Head crossover on the same L200 candidates and features. Recall is official query-macro [email protected] at the indicated global return budget. Each head’s tF1 uses its own global test-F1-optimal threshold; these are diagnostic operating points. Recall and tF1 are percentages.
AdamW; generator LR 2×10−5 , head LR 2×10−4 ; cosine decay, 3% warmup
SFT schedule
60,405 units; global batch 64; two epochs, 1,888 steps
RL actor
Full-parameter updates; 200 outer steps; LR 10−6 for steps 1–100, then 2×10−7
Appendix
Table 10: Training settings for the main-table models.
Figure 4: Ranking: a workflow needs one useful clip from several proposals. (a) Confidence selects the later washing event instead of the first generated interval. (b) For replacing the bottle cap, it selects an existing boundary alternative with tIoU 0.905; the earlier C2 has tIoU 0.600. This is selection among decoded intervals, with no boundary refinement. All raw candidates are shown in generation order; the highest score supplies the single answer, without NMS or thresholding. Across the case figures, teal is GT, purple is selected, and gray is unselected. Scores are rounded only for display; all times are clip-relative seconds.
Figure 5: Set filtering: review repeated events with fewer low-quality intervals. Both OMTG-Bench cases use generation-order NMS at tIoU 0.3 followed by that variant’s global test-oracle threshold, τ≈0.318 . Every NMS survivor is shown, with its original candidate index. (a) A misplaced interval survives NMS but is removed by confidence, retaining both annotated occurrences. (b) Four workout intervals are retained while the extra [114,129] fragment is removed. F1 rises from 0.800 and 0.889 to 1.000, respectively; recall remains 1.000 in both cases. The same threshold supports different answer counts.
Figure 6: Query-level rejection: a retrieval interface may need to return no result. Each video is paired with its original matched query and one synthetic cross-video mismatch from the accepted mixed-query evaluation. Only the highest-confidence raw candidate for each query is displayed; the generator produced nonempty candidates for all four queries. For this raw-pool diagnostic, the benchmark-wide test-oracle thresholds are 0.549 for Charades and 0.524 for QVHighlights, with no per-case tuning and no NMS. Both matched queries are accepted, while both mismatches return [] . An asterisk denotes a synthetic unmatched label, not an official exhaustively verified negative annotation.
Figure 7: Two routes to a missed occurrence. Selected OMTG-Bench cases from the final L200 model with 150 head updates. (a) A close-up of the second annotation in Q58: generation-order NMS suppresses C4 despite its higher confidence and GT IoU. (b) All six raw candidates in Q149 miss G1; purple rows are retained by NMS and the model’s global F1 test-oracle threshold. Matrix entries are original-coordinate tIoUs; matching uses tIoU ≥0.5 . Times are clip-relative seconds. Frames provide scene context, and candidate IDs preserve generation order.
Media Analytics and Computing Lab, Xiamen University, Xiamen, China · Kling Team, Kuaishou Technology, China · The Chinese University of Hong Kong, HKSAR, China +1