Existing multimodal video highlight detectors typically assume that visual, audio, and textual streams are continuously available. In practice, however, inputs may suffer from localized frame missingness or complete-stream outage. We formulate this robustness challenge along two dimensions: temporal missingness, where frames are missing independently in each modality, and stream-level missingness, where one modality is unavailable throughout a video. Moreover, we find that the mean squared error (MSE) loss is misaligned with both the evaluation metrics and the peak-driven nature of highlights. Therefore, we propose Temporal-Stream Modality Dropout (TSMD), which combines structured missingness simulation with a joint objective comprising pointwise MSE, per-video Pearson correlation, and peak-oriented RankNet loss terms. TSMD has three variants: temporal, stream-level, and mixed dropout. On the MoSu and Mr. HiSum datasets, TSMD-Temporal improves mAP@15 by 7.06 and 3.41 points over TripleSumm under 50% independent temporal removal, whereas TSMD-Stream performs the best under complete-stream removal. TSMD-Mix retains most of these complementary benefits and ranks the best or the second-best across the evaluated temporal and stream-level conditions.
Figures & tables
Figure 1: Incomplete multimodal inputs in video highlight detection. Left: performance degrades under asynchronous V isual, A udio, and T ext corruption in deployment. Right: our retrained TripleSumm results on MoSu (Table 1 ). The missing-input panel reports 50% independent temporal removal and mean complete-stream removal, with clean performance shown as a dashed reference.
Figure 2: Multimodal highlight detection and TSMD masking patterns: (a) independent temporal and (b) complete-stream dropout, illustrated for the visual stream.
MoSu
Mr. HiSum
Method / training policy
Mod.
τ↑
ρ↑
mAP@50 ↑
mAP@15 ↑
τ↑
ρ↑
mAP@50 ↑
mAP@15 ↑
(a) Clean inputs
VASNet [ 3 ]
V
0.151
0.219
64.49
31.05
0.069
0.102
58.69
25.28
A2Summ [ 6 ]
VT
0.181
0.257
66.48
35.70
0.121
0.172
63.20
32.34
UMT [ 10 ]
VA
0.239
0.334
68.83
36.73
0.178
0.253
66.81
35.65
CSTA [ 16 ]
V
0.291
0.398
71.77
40.65
0.128
0.185
63.38
30.42
Table 1: Results on MoSu and Mr. HiSum datasets under (a) clean inputs, (b) independent temporal removal ( r=0.5 ), and (c) complete-stream removal (averaged across − V/ − A/ − T). Published baselines from [ 9 ] appear above the dashed line in (a). All other rows are our three-seed mean ± SD. TripleSumm † denotes our retrained MSE baseline, and all TSMD variants use the joint objective. The best and second best are marked per column; τ and ρ are ranked at full precision. Mod.: input modalities.
Objective
MoSu
Mr. HiSum
τ
ρ
mAP@15
τ
ρ
mAP@15
M
0.353
0.473
44.81
0.255
0.347
41.12
M+P
0.361
0.481
45.59
0.263
0.356
41.64
M+R
0.353
0.475
45.82
0.249
0.341
41.81
M+P+R
0.360
0.481
45.65
0.260
0.353
42.31
Table 2: Three-seed mean results for objective components under clean inputs, where M, P, and R denote MSE, Pearson, and RankNet losses, respectively.
Dataset
Training policy
−V
−A
−T
MoSu
Clean training
36.37
38.99
42.60
+ TSMD-Stream
42.07
43.37
44.46
Mr. HiSum
Clean training
38.90
36.77
41.77
+ TSMD-Stream
41.03
38.62
41.55
Table 3: Three-seed mean test mAP@15 under each missing stream using the joint objective.
Figure 3: MoSu test mAP@15 under increasing independent temporal removal ratio. Blue and red curves use the MSE-only and joint objectives. Solid and dashed curves denote clean and temporal-dropout training, respectively. Their separation shows the complementary contributions of the joint objective and temporal-dropout training.
Detecting video highlights, the most informative or engaging moments in a video, is important for applications such as video summarization and content recommendation. We propose a zero-shot framework that combines CLIP, large language models (LLMs), and diffusion models. Given lightweight video metadata, such as a title or category, an LLM generates textual descriptions of likely highlight events. These descriptions are further converted into synthetic visual prototypes using a diffusion model. Textual and visual representations are matched to video frames using CLIP, enabling frame-level highlight detection without highlight annotations or dataset-specific training. Experiments on TVSum and SumMe demonstrate strong zero-shot performance, with particularly favorable results on TVSum. The proposed approach provides an effective framework for metadata-conditioned zero-shot video highlight detection.
Michal Byra, Alberto Presta, Grzegorz Stefanski +1
Samsung AI Center, Warsaw, Poland · Institute of Fundamental Technological Research, Polish Academy of Sciences, Warsaw, Poland
Video Moment Retrieval (MR) and Highlight Detection (HD) are crucial tasks in video analysis that aim to localize specific moments and estimate clip-wise relevance based on a given text query. Recent approaches treat them as similar video grounding tasks and use the same architecture to solve them. These tasks require both fine-grained comprehension at the image level and high-level temporal understanding across the entire video. Existing approaches have primarily focused on temporal modeling using frame-level features, often neglecting the rich visual information related to the text query within individual frames. This oversight leads to inaccurate grounding results. To address this limitation, we propose a Comprehensive Spatial-Temporal Representation Learning Framework (CoSTL), which captures both fine-grained image-level information and temporal dynamics. Specifically, CoSTL incorporates a text-driven progressive fine-grained image encoder, performing a two-step text-driven knowledge extraction process to learn fine-grained spatial representations. Furthermore, a multi-scale temporal perception module captures comprehensive spatial-temporal representations, enhancing the model's ability to process temporal dynamics. We demonstrate state-of-the-art performance on four public benchmarks: QVHighlights, Charades-STA, TACoS, and TVSum.
Xin Dong, Wenjia Geng, Wenfeng Deng +1
Shenzhen International Graduate School, Tsinghua University · Pengcheng Laboratory
Video understanding has become more and more important with the growth of Artificial Intelligence (AI) for video generation. Recently, Multimodal Large Language Model(M-LLM) has shown its capability in video understanding. Video summarization, a specific domain of video understanding, has proven its importance for efficient navigation and retrieval. Both video understanding and video summarization require a good selection of key frames in a video. Current video summarization methods heavily focus on the selected key frames and correlated segment captions. However, existing approaches overlook the perspective of treating the importance of the frames globally. We argue that using discrete selected frames for summarization will not only reduce the understanding coherence, but also lost important information in the video, as well as wasting the original capacity of the MLLMs. In this paper, we propose HAS, a Highlight-guided Attention Steering method for video summarization. We consider a challenging but practical setting where the video given to MLLMs for summarize should be continuous but with highlight guidance. HAS mainly consists of two parts: The first part is to find a continuous frame-level highlight distribution for the video globally. The second part is to apply the highlight distribution as an attention steering vector for the MLLM, targeting a better understanding of the video, and thus during the model inference time, putting more attention on the highlighted frames, while avoiding lost entire information on less highlighted frames through putting less attention instead of forgetting them. We evaluated HAS on a variety of benchmarks, and it has shown convincing performance in video summarization.