cs.CVMay 30, 2025

LLM-powered Query Expansion for Enhancing Boundary Prediction in Language-driven Action Localization

Authors: Zirui Shang, Xinxiao Wu, Shuo Yang

Organizations: Beijing Key Laboratory of Intelligent Information Technology, School of Computer Science & Technology, Beijing Institute of Technology, China. · Guangdong Laboratory of Machine Perception and Intelligent Computing, Shenzhen MSU-BIT University, China.

Abstract

Language-driven action localization in videos requires not only semantic alignment between language query and video segment, but also prediction of action boundaries. However, the language query primarily describes the main content of an action and usually lacks specific details of action start and end boundaries, which increases the subjectivity of manual boundary annotation and leads to boundary uncertainty in training data. In this paper, on one hand, we propose to expand the original query by generating textual descriptions of the action start and end boundaries through LLMs, which can provide more detailed boundary cues for localization and thus reduce the impact of boundary uncertainty. On the other hand, to enhance the tolerance to boundary uncertainty during training, we propose to model probability scores of action boundaries by calculating the semantic similarities between frames and the expanded query as well as the temporal distances between frames and the annotated boundary frames. They can provide more consistent boundary supervision, thus improving the stability of training. Our method is model-agnostic and can be seamlessly and easily integrated into any existing models of language-driven action localization in an off-the-shelf manner. Experimental results on several datasets demonstrate the effectiveness of our method.

Figures & tables

Explore similar work

CardsList
  1. Masked Diffusion Vision-Language Models for Temporal Action Localization

    May 28, 2026Fengshun Wang, Zhengbo Zhang, Zhigang TuHuman Activity RecognitionVideo-Language Models

  2. Zero-Shot Temporal Action Localization Through Textual Guidance

    May 21, 2026Benedetta Liberatori, Alessandro Conti, Lorenzo Vaquero +3Zero-Shot LearningTemporal Action Detection

  3. Language-Augmented Video Action Anticipation: Design Fundamentals, Benchmarks, and Open Challenges

    Sep 15, 2026Mahsa Mohammadi, Zeyu Fu, Sareh Rowlands