Existing multimodal video highlight detectors typically assume that visual, audio, and textual streams are continuously available. In practice, however, inputs may suffer from localized frame missingness or complete-stream outage. We formulate this robustness challenge along two dimensions: temporal missingness, where frames are missing independently in each modality, and stream-level missingness, where one modality is unavailable throughout a video. Moreover, we find that the mean squared error (MSE) loss is misaligned with both the evaluation metrics and the peak-driven nature of highlights. Therefore, we propose Temporal-Stream Modality Dropout (TSMD), which combines structured missingness simulation with a joint objective comprising pointwise MSE, per-video Pearson correlation, and peak-oriented RankNet loss terms. TSMD has three variants: temporal, stream-level, and mixed dropout. On the MoSu and Mr. HiSum datasets, TSMD-Temporal improves mAP@15 by 7.06 and 3.41 points over TripleSumm under 50% independent temporal removal, whereas TSMD-Stream performs the best under complete-stream removal. TSMD-Mix retains most of these complementary benefits and ranks the best or the second-best across the evaluated temporal and stream-level conditions.
Figures & tables
Figure 1: Incomplete multimodal inputs in video highlight detection. Left: performance degrades under asynchronous V isual, A udio, and T ext corruption in deployment. Right: our retrained TripleSumm results on MoSu (Table 1 ). The missing-input panel reports 50% independent temporal removal and mean complete-stream removal, with clean performance shown as a dashed reference.
Figure 2: Multimodal highlight detection and TSMD masking patterns: (a) independent temporal and (b) complete-stream dropout, illustrated for the visual stream.
MoSu
Mr. HiSum
Method / training policy
Mod.
τ↑
ρ↑
mAP@50 ↑
mAP@15 ↑
τ↑
ρ↑
mAP@50 ↑
mAP@15 ↑
(a) Clean inputs
VASNet [ 3 ]
V
0.151
0.219
64.49
31.05
0.069
0.102
58.69
25.28
A2Summ [ 6 ]
VT
0.181
0.257
66.48
35.70
0.121
0.172
63.20
32.34
UMT [ 10 ]
VA
0.239
0.334
68.83
36.73
0.178
0.253
66.81
35.65
CSTA [ 16 ]
V
0.291
0.398
71.77
40.65
0.128
0.185
63.38
30.42
Table 1: Results on MoSu and Mr. HiSum datasets under (a) clean inputs, (b) independent temporal removal ( r=0.5 ), and (c) complete-stream removal (averaged across − V/ − A/ − T). Published baselines from [ 9 ] appear above the dashed line in (a). All other rows are our three-seed mean ± SD. TripleSumm † denotes our retrained MSE baseline, and all TSMD variants use the joint objective. The best and second best are marked per column; τ and ρ are ranked at full precision. Mod.: input modalities.
Objective
MoSu
Mr. HiSum
τ
ρ
mAP@15
τ
ρ
mAP@15
M
0.353
0.473
44.81
0.255
0.347
41.12
M+P
0.361
0.481
45.59
0.263
0.356
41.64
M+R
0.353
0.475
45.82
0.249
0.341
41.81
M+P+R
0.360
0.481
45.65
0.260
0.353
42.31
Table 2: Three-seed mean results for objective components under clean inputs, where M, P, and R denote MSE, Pearson, and RankNet losses, respectively.
Dataset
Training policy
−V
−A
−T
MoSu
Clean training
36.37
38.99
42.60
+ TSMD-Stream
42.07
43.37
44.46
Mr. HiSum
Clean training
38.90
36.77
41.77
+ TSMD-Stream
41.03
38.62
41.55
Table 3: Three-seed mean test mAP@15 under each missing stream using the joint objective.
Figure 3: MoSu test mAP@15 under increasing independent temporal removal ratio. Blue and red curves use the MSE-only and joint objectives. Solid and dashed curves denote clean and temporal-dropout training, respectively. Their separation shows the complementary contributions of the joint objective and temporal-dropout training.