cs.CVAug 21, 2025

Diffusion-grounded VideoLLM for Entity-aware temoporal grounding

Authors: Pengcheng Fang, Yuxia Chen, Yong Wang, Xiaohao Cai, Xiangning Ruan

Abstract

Precise temporal grounding requires distinguishing when a queried event occurs from when its participating entities are merely visible. We propose Diffusion-Grounded VideoLLM, which conditions temporal feature extraction on query-relevant entities before language reasoning. The framework tracks entities named in the query and uses their masks to condition a frozen video diffusion backbone. Intermediate spatiotemporal features are extracted through truncated denoising and combined with entity tokens and timestamp embeddings. The language model uses this evidence together with the full query to generate temporal intervals and answers to grounded questions. On Charades-STA and NExT-GQA, the model obtains 43.5 mIoU and 28.4 Acc@GQA, improving the reported Grounded-VideoLLM reference by 6.7 and 1.7 points, respectively. Component and entity-pathway ablations support the usefulness of conditioning diffusion features on query-relevant entities for temporal grounding.

Figures & tables

Explore similar work

CardsList
  1. IAE-VTG: Interaction-Aligned Action-Entity Video Temporal Grounding

    Sep 9, 2026Shiwen Zhao, Qi Zhang, Sezer Karaoglu +2Video Temporal GroundingVideo Understanding

  2. Foresee-to-Ground: From Predictive Temporal Perception to Evidence-Driven Reasoning for Video Temporal Grounding

    May 21, 2026Zelin Zheng, Xinyan Liu, Ruixin Li +4Video Temporal GroundingTemporal Grounding