cs.MMJan 14, 2025

Zero-shot Video Moment Retrieval via Off-the-shelf Multimodal Large Language Models

Authors: Yifang Xu, Yunzhuo Sun, Benxiang Zhai, Ming Li, Wenxin Liang, Yang Li, Sidan Du

Organizations: Nanjing University · Dalian University of Technology · Nanjing University of Information Science and Technology

Abstract

The target of video moment retrieval (VMR) is predicting temporal spans within a video that semantically match a given linguistic query. Existing VMR methods based on multimodal large language models (MLLMs) overly rely on expensive high-quality datasets and time-consuming fine-tuning. Although some recent studies introduce a zero-shot setting to avoid fine-tuning, they overlook inherent language bias in the query, leading to erroneous localization. To tackle the aforementioned challenges, this paper proposes Moment-GPT, a tuning-free pipeline for zero-shot VMR utilizing frozen MLLMs. Specifically, we first employ LLaMA-3 to correct and rephrase the query to mitigate language bias. Subsequently, we design a span generator combined with MiniGPT-v2 to produce candidate spans adaptively. Finally, to leverage the video comprehension capabilities of MLLMs, we apply VideoChatGPT and span scorer to select the most appropriate spans. Our proposed method substantially outperforms the state-ofthe-art MLLM-based and zero-shot models on several public datasets, including QVHighlights, ActivityNet-Captions, and Charades-STA.

Explore similar work

CardsList
  1. Mitigating Modality and Language-Style Gaps for Zero-Shot Video Moment Retrieval

    Jul 21, 2026Jihyun Lee, Cheol-Ho Cho, Woojin Jun +2Zero-ShotModalities

  2. Fewer Steps, Better Performance: Efficient Cross-Modal Clip Trimming for Video Moment Retrieval Using Language

    May 28, 2026Xiang Fang, Daizong Liu, Wanlong Fang +5

  3. Retrieving Any Relevant Moments: Benchmark and Models for Generalized Moment Retrieval

    May 4, 2026Yiming Ding, Siyu Cao, Luyuan Jiao +4Multimodal RetrievalObject Localization