cs.CVSep 30, 2026

Grounding with Confidence: Controllable Generative Video Temporal Grounding

Authors: Jinhao Chen, Benlei Cui, Ruijian Jia, Ziheng Wang, Tianyu Wo, Pengfei Sun, Longtao Huang, Hui Xue, +2 more

Organizations: Alibaba Group · Beihang University · Fudan University

Abstract

Video temporal grounding supports applications such as video search, content review, and automated editing by localizing events described in natural language. Yet existing generative models typically output timestamps without explicit interval-level confidence scores to guide candidate selection. We separate candidate generation from acceptance by scoring individual intervals within the original decoding pass. A lightweight confidence head reads pooled decoder states, providing an explicit score trained for interval selection. Offline verifier scores supervise the head on fixed candidate sequences, and temporal-overlap labels adapt it to current rollouts during reinforcement learning. GT-anchored candidate-pool supervision and set-level optimization train the generator. The resulting scores support ranking, threshold-based selection, and rejection without invoking an external verifier at inference. On a fixed OMTG-Bench candidate pool, confidence raises query-macro [email protected] from 9.95% to 14.42% over generation order at a 10% global return budget, and from 26.48% to 31.12% at a 25% budget. The continuous scores let downstream applications adjust return budgets or acceptance thresholds to match their precision-recall preferences, without regenerating candidate intervals.

Figures & tables

Appendix figures & tables10 assets

Supplementary material from the paper’s appendix.

Appendix

Explore similar work

CardsList
  1. REZE: Recognition-Based Zero-Shot Extraction for Video Temporal Grounding

    Aug 5, 2026Boyang Li, Chenhui Gou, Jianfei CaiTemporal Video GroundingZero-Shot Learning

  2. TimeLens2: Generalist Video Temporal Grounding with Multimodal LLMs

    Jul 19, 2026Yuhan Zhu, Changlian Ma, Xiangyu Zeng +12Temporal Video GroundingReinforcement Learning

  3. TimePLE: Rethinking Temporal Representation for Video Temporal Grounding

    Jul 27, 2026Yuhui Zeng, Xinyu Mao, Xiaokun Liu +4Temporal Video GroundingRepresentation Learning