Grounding with Confidence: Controllable Generative Video Temporal Grounding
Organizations: Alibaba Group · Beihang University · Fudan University
Abstract
Video temporal grounding supports applications such as video search, content review, and automated editing by localizing events described in natural language. Yet existing generative models typically output timestamps without explicit interval-level confidence scores to guide candidate selection. We separate candidate generation from acceptance by scoring individual intervals within the original decoding pass. A lightweight confidence head reads pooled decoder states, providing an explicit score trained for interval selection. Offline verifier scores supervise the head on fixed candidate sequences, and temporal-overlap labels adapt it to current rollouts during reinforcement learning. GT-anchored candidate-pool supervision and set-level optimization train the generator. The resulting scores support ranking, threshold-based selection, and rejection without invoking an external verifier at inference. On a fixed OMTG-Bench candidate pool, confidence raises query-macro [email protected] from 9.95% to 14.42% over generation order at a 10% global return budget, and from 26.48% to 31.12% at a 25% budget. The continuous scores let downstream applications adjust return budgets or acceptance thresholds to match their precision-recall preferences, without regenerating candidate intervals.
Figures & tables
| Budget | Order | Logp | Verifier | Confidence |
|---|---|---|---|---|
| 10% | 9.95 | 9.41 | 11.00 | 14.42 |
| 25% | 26.48 | 22.37 | 26.31 | 31.12 |
| 50% | 51.34 | 43.05 | 47.72 | 52.82 |
| 75% | 64.76 | 61.00 | 62.45 | 65.30 |
| Model | C-Acc | [email protected] | [email protected] | tIoU | EtF1 |
|---|---|---|---|---|---|
| Seed-1.8 | 38.12 | 67.13 | 54.67 | 56.81 | 28.04 |
| Gemini 2.5 Pro | 50.94 | 55.72 | 43.57 | 43.24 | 27.80 |
| Qwen3-VL-4B | 0.31 | 37.07 | 26.75 | 30.42 | 0.21 |
| TimeLens-8B | 0.00 | 39.14 | 32.76 | 32.38 | 0.00 |
| TimeLens2-4B | 18.75 | 53.02 | 43.89 | 48.73 | 14.29 |
| TimeLens2-8B | 16.25 | 47.14 | 38.19 | 47.08 | 12.86 |
| Raw | NMS | NMS + confidence † | ||||
|---|---|---|---|---|---|---|
| Model | Recall | tF1 | Recall | tF1 | Recall | tF1 |
| SFT | 67.85 | 34.65 | 65.06 | 64.17 | 64.92 | 64.54 |
| RL | 73.64 | 41.39 | 70.51 | 67.07 | 70.10 | 67.39 |
| Charades | ActivityNet | QVHighlights | ||||
|---|---|---|---|---|---|---|
| Method | [email protected] | FAR | [email protected] | FAR | [email protected] | FAR |
| TimeLens2-4B, native | 44.24 | 91.17 | 47.38 | 66.04 | 54.95 | 71.64 |
| TimeLens2-8B, native | 44.87 | 88.25 | 44.97 | 78.13 | 52.57 | 88.38 |
| Ours (SFT) † | 46.81 | 36.01 | 48.42 | 15.22 | 57.67 | 16.87 |
| Ours (RL) † | 49.87 | 21.35 | 48.78 | 13.24 | 56.52 | 18.95 |
Appendix figures & tables10 assets
Supplementary material from the paper’s appendix.
Appendix
| Budget | vs. Order | vs. Logp | vs. Verifier |
|---|---|---|---|
| 10% | +4.47 [2.30, 6.04] | +5.01 [3.26, 6.36] | +3.42 [1.36, 4.56] |
| 25% | +4.64 [2.26, 7.09] | +8.75 [5.86, 11.29] | +4.81 [2.33, 7.16] |
| mIoU | R@1, 0.5 | ||||
| Dataset | Queries | First | Confidence | mIoU [95% CI] | First Conf. |
| Charades | 3,363 | 47.89 | 49.91 | +2.02 [1.39, 2.68] | 55.16 57.21 |
| ActivityNet | 4,500 | 45.96 | 46.40 | +0.44 [ 0.21, 1.11] | 50.76 51.33 |
| QVHighlights | 1,541 | 57.26 | 57.95 | +0.68 [ 0.38, 1.69] | 61.19 62.10 |
| Model | Selection | [email protected] | [email protected] | [email protected] |
|---|---|---|---|---|
| SFT | NMS | 66.17 | 65.06 | 64.17 |
| SFT | NMS + confidence | 67.04 | 64.92 | 64.54 |
| RL | NMS | 66.25 | 70.51 | 67.07 |
| RL | NMS + confidence | 67.23 | 70.10 | 67.39 |
| Variant | Recall @0.5 | C-Acc | tF1 @0.5 | tIoU | EtF1 |
|---|---|---|---|---|---|
| Candidate generation: raw outputs | |||||
| GT-only SFT | 61.55 | 48.13 | 62.59 | 58.67 | 38.08 |
| Ours (SFT), raw | 67.85 | 0.00 | 34.65 | 53.35 | 0.00 |
| Ours (RL), raw | 73.64 | 0.31 | 41.39 | 57.21 | 0.01 |
| Final selection: NMS + confidence selection | |||||
| Ours (SFT) † | 64.92 | 51.88 | 64.54 | 60.75 | 40.26 |
| Head | NMS AUROC | Oracle [email protected] | Recall, 10% | Recall, 25% |
|---|---|---|---|---|
| Original SFT | 0.754 | 67.18 | 12.80 | 29.74 |
| Online RL | 0.818 | 67.39 | 14.42 | 31.12 |
| Component | Setting |
|---|---|
| Base model | TimeLens2-4B; hidden size 2,560 |
| Confidence head | LayerNorm, Linear(2560, 256), GELU, dropout 0.1, Linear(256, 1), sigmoid |
| SFT adaptation | LoRA rank 32, alpha 64, dropout 0.05; fresh head |
| SFT optimizer | AdamW; generator LR , head LR ; cosine decay, 3% warmup |
| SFT schedule | 60,405 units; global batch 64; two epochs, 1,888 steps |
| RL actor | Full-parameter updates; 200 outer steps; LR for steps 1–100, then |