LLM-powered Query Expansion for Enhancing Boundary Prediction in Language-driven Action Localization
Authors: Zirui Shang, Xinxiao Wu, Shuo Yang
Organizations: Beijing Key Laboratory of Intelligent Information Technology, School of Computer Science & Technology, Beijing Institute of Technology, China. · Guangdong Laboratory of Machine Perception and Intelligent Computing, Shenzhen MSU-BIT University, China.
Language-driven action localization in videos requires not only semantic alignment between language query and video segment, but also prediction of action boundaries. However, the language query primarily describes the main content of an action and usually lacks specific details of action start and end boundaries, which increases the subjectivity of manual boundary annotation and leads to boundary uncertainty in training data. In this paper, on one hand, we propose to expand the original query by generating textual descriptions of the action start and end boundaries through LLMs, which can provide more detailed boundary cues for localization and thus reduce the impact of boundary uncertainty. On the other hand, to enhance the tolerance to boundary uncertainty during training, we propose to model probability scores of action boundaries by calculating the semantic similarities between frames and the expanded query as well as the temporal distances between frames and the annotated boundary frames. They can provide more consistent boundary supervision, thus improving the stability of training. Our method is model-agnostic and can be seamlessly and easily integrated into any existing models of language-driven action localization in an off-the-shelf manner. Experimental results on several datasets demonstrate the effectiveness of our method.
Figures & tables
Figure 1: Illustration of boundary uncertainty for the same language query. The orange text represents the motion of the start boundary annotation and the blue text represents the motion of the end boundary annotation.
Figure 2: Overview of the proposed action localization framework. The pipeline first leverages a large language model to perform text expansion, generating detailed sentences that describe the beginning and ending stages of the target action based on the original content-focused query. These expanded boundary cues, along with the visual video, are processed by the base model to obtain initial text and video features. To refine the multimodal representations, we introduce a query-guided temporal modeling module that fuses the detailed start and end descriptions with the video features, employing global and local branches to capture holistic context and stage-specific boundary details. Finally, a boundary probability modeling module is proposed to estimate the likelihood of action boundaries across the temporal sequence, converting rigid ground-truth intervals into soft probability scores for more robust supervision during training.
Figure 3: An example of the designed question and the corresponding answers from LLM.
Table 2: Results of different proposed modules with QD-DETR across three datasets. “QTM” and “BPM” denote Query-guided Temporal Modeling and Boundary Probability Modeling, respectively.
Table 7: Performance comparison on different action location subsets on Charades-STA.
Figure 8: Quantitative evaluation of query expansion quality by GPT-4 and human raters.
Figure 9: Examples of action localization results with similar language queries on the Charades-STA dataset by integrating our modules into QD-DETR. The texts in red are the expanded start queries, and the texts in blue are the expanded end queries.
Figure 10: Visualization of boundary probability scores on the Charades-STA dataset by integrating our modules into QD-DETR.
Process
Time (ms)
FLOPs (G)
Query Expansion
∼ 1315
1442.7
Visual Encoder
∼ 1344
54.9
Text Encoder
∼ 2.90
1.38
QTM
∼ 19.83
3.09
Localization Head
∼ 17.87
2.46
Table 8: Inference latency and computational cost (per sample) on the TACoS dataset.