Advancing Video-Text Pretraining with Multi-View Captions
Organizations: King Abdullah University of Science and Technology (KAUST) · University of Li`ege
Abstract
Video-text pretraining has achieved remarkable progress through the scaling of models and datasets, yet the quality of language supervision remains underexplored. Existing web-scale datasets often provide only a single sparse caption per video that fails to capture rich spatiotemporal semantics, while directly using captioning models can generate noisy descriptions. We propose a large-scale multimodal large language model-based supervision generation framework that improves supervision diversity, fidelity, and semantic coverage. Starting from 10 million videos, our approach generates multi-view captions (MVC) through complementary summary and detailed captions, reasoning-based refinement, and semantic positive caption generation. To effectively exploit supervision at different granularities, we further introduce a granularity-aware text representation with separate CLS tokens for summary and detailed views. We pretrain video-text models using the resulting supervision corpus and evaluate them across standard, fine-grained and detailed text-to-video retrieval benchmarks. Our approach consistently improves both zero-shot and fine-tuned performance while using smaller pretraining corpora than existing methods, demonstrating the importance of rich and complementary textual supervision for video-text pretraining. Project page: https://rvandeghen.github.io/mvc/
Figures & tables
| Method | #Samples | Source | MSR-VTT | DiDeMo | ActivityNet | LSMDC | MSVD |
|---|---|---|---|---|---|---|---|
| R@1 | R@1 | R@1 | R@1 | R@1 | |||
| Frozen ( Bain et al., 2021 ) | 5M | WebVid-2M+Img3M | 18.7 | 20.2 | – | – | – |
| Singularity ( Lei et al., 2023 ) | 5M | WebVid-2M+Img3M | 28.4 | 36.9 | 30.8 | – | – |
| UMT-B ( Li et al., 2023b ) | 5M | WebVid-2M+Img3M | 29.6 | 33.4 | 28.3 | 16.8 | 36.2 |
| STM-B ( Wu et al., 2025 ) | 5M | WebVid-2M+Img3M | 29.8 | 34.4 | 30.6 | – | 38.7 |
| ClusterSTM-B ( Zhuang et al., 2026 ) | 5M | WebVid-2M+Img3M | 31.2 | 36.5 | 31.4 | – | 40.3 |
| Method | #Samples | Source | MSR-VTT | DiDeMo | ActivityNet | LSMDC | MSVD |
|---|---|---|---|---|---|---|---|
| R@1 | R@1 | R@1 | R@1 | R@1 | |||
| ClipBERT ( Lei et al., 2021 ) | 5.4M | COCO+VG | 22.0 | 20.4 | 21.3 | – | – |
| Frozen ( Bain et al., 2021 ) | 5M | WebVid-2M+Img3M | 31.0 | 34.6 | – | 15.0 | 33.7 |
| UMT-B ( Li et al., 2023b ) | 5M | WebVid-2M+Img3M | 46.3 | 54.8 | 52.1 | 30.3 | 47.4 |
| STM-B ( Wu et al., 2025 ) | 5M | WebVid-2M+Img3M | 48.5 | 56.9 | 53.6 | – | – |
| ClusterSTM-B ( Zhuang et al., 2026 ) | 5M | WebVid-2M+Img3M | 49.7 | 58.5 | 54.9 | – | – |
| Method | #Samples | Source | CARE-S | CARE-T | DREAM-D | DREAM-E | S2S-W | S2S-S |
| R@1 | R@1 | R@1 | R@1 | R@1 | R@1 | |||
| CLIP-B/16 ( Radford et al., 2021 ) | 400M | CLIP-400M | 45.6 | 30.3 | 32.6 | 13.6 | 10.7 | 22.7 |
| CLIP-L/14 ( Radford et al., 2021 ) | 400M | CLIP-400M | 49.1 | 33.5 | 44.3 | 14.6 | 65.8 | 45.4 |
| ViCLIP-B ( Wang et al., 2024b ) | 400M+ 10M | CLIP-400M + IV-10M-FLT | 56.0 | 31.3 | 54.9 | 10.4 | 52.2 | 43.9 |
| ViCLIP-L ( Wang et al., 2024b ) | 400M + 10M | CLIP-400M + IV-10M-FLT | 55.7 | 34.5 | 55.5 | 23.3 | 55.1 | 43.6 |
| Long-CLIP-L/14 ( Zhang et al., 2024 ) | 400M | CLIP-400M + SG4V-1M | 65.6 | 33.3 | 58.3 | 24.3 | 74.7 | 47.7 |
| Supervision | MSVD | ANet | DiDeMo | LSMDC | CARE-T | Shot2Story-S |
|---|---|---|---|---|---|---|
| Original only ( ) | 35.4 | 25.9 | 29.3 | 11.8 | 27.1 | 42.2 |
| Ours: summary only ( ) | 37.0 | 46.2 | 46.0 | 15.1 | 48.0 | 63.2 |
| Ours: full ( ) | 38.5 | 50.1 | 52.4 | 16.7 | 54.1 | 67.4 |
| Original + ours ( ) | 38.8 | 50.1 | 53.0 | 18.2 | 54.1 | 67.9 |
| Supervision | MSVD | ANet | DiDeMo | LSMDC | CARE-T | Shot2Story-S |
|---|---|---|---|---|---|---|
| Summary only ( ) | 37.0 | 46.2 | 46.0 | 15.1 | 48.0 | 63.2 |
| Detailed only ( ) | 26.4 | 42.7 | 38.4 | 12.4 | 41.1 | 62.3 |
| Summary+detailed ( ) | 36.6 | 48.5 | 48.6 | 15.9 | 50.1 | 65.9 |
| Summary+refined detailed ( ) | 37.5 | 49.6 | 49.7 | 16.1 | 51.8 | 66.7 |
| Summary+refined+positive ( ) | 38.5 | 50.1 | 52.4 | 16.7 | 54.1 | 67.4 |
| Text Representation | MSVD | ANet | DiDeMo | LSMDC | CARE-T | Shot2Story-S |
|---|---|---|---|---|---|---|
| Single CLS | 38.8 | 50.1 | 53.0 | 18.2 | 54.1 | 67.9 |
| Dual CLS (Ours) | 39.4 | 51.3 | 54.3 | 18.7 | 56.4 | 68.5 |
Appendix figures & tables7 assets
Supplementary material from the paper’s appendix.
Appendix
| Method | #Samples | ANet-QA | MSR-VTT-QA | MSVD-QA |
|---|---|---|---|---|
| ClipBERT ( Lei et al., 2021 ) | 0.2M | – | 37.4 | – |
| ALPRO ( Li et al., 2022a ) | 5M | – | 42.1 | 45.9 |
| JustAsk ( Yang et al., 2021 ) | 69M | 38.9 | 41.5 | 47.5 |
| VideoCLIP ( Xu et al., 2021 ) | 136M | – | – | – |
| All-in-one ( Wang et al., 2023 ) | 138M | – | 44.3 | 47.9 |
| MERLOT ( Zellers et al., 2021 ) | 180M | 41.4 | 43.1 | – |
| Caption type | Avg. words | Std. words |
|---|---|---|
| Original caption | 9.4 | 2.6 |
| Qwen summary caption | 13.8 | 2.6 |
| Tarsier summary caption | 16.2 | 4.5 |
| Qwen detailed caption | 96.0 | 27.1 |
| Tarsier detailed caption | 80.1 | 25.0 |
| Qwen refined detailed caption | 60.6 | 19.5 |
| Masking ratio | MSVD | ANet | DiDeMo | LSMDC | CARE-T | Shot2Story-S |
|---|---|---|---|---|---|---|
| 50% | 38.8 | 50.1 | 53.0 | 18.2 | 54.1 | 67.9 |
| 25% | 39.5 | 51.9 | 53.0 | 15.8 | 53.3 | 68.5 |
| 10% | 39.7 | 52.5 | 54.3 | 15.7 | 55.9 | 68.3 |
| Epochs | MSVD | ANet | DiDeMo | LSMDC | CARE-T | Shot2Story-S |
|---|---|---|---|---|---|---|
| 5 | 31.6 | 40.8 | 42.6 | 12.9 | 42.9 | 59.7 |
| 10 | 36.2 | 46.5 | 48.7 | 15.8 | 50.3 | 64.9 |
| 15 | 37.9 | 48.9 | 51.5 | 17.3 | 52.6 | 66.7 |
| 20 | 38.8 | 50.1 | 53.0 | 18.2 | 54.1 | 67.9 |
| Task | Dataset | LR | Epochs | DropPath |
|---|---|---|---|---|
| Retrieval | MSR-VTT | 2e-5 (B), 1e-5 (L) | 10(B),7(L) | 0.2(B),0.3(L) |
| DiDeMo | 2e-5(B), 2e-5(L) | 12(B),5(L) | 0.1(B),0.3(L) | |
| ActivityNet | 4e-5(B), 1e-5 (L) | 20(B/L) | 0.1(B),0.2(L) | |
| LSMDC | 2e-5 (B), 1e-5 (L) | 10(B),8(L) | 0.1(B),0.2(L) | |
| MSVD | 2e-5 (B), 1e-5 (L) | 10(B/L) | 0.2(B),0.3(L) | |
| VQA | ActivityNet-QA | 4e-5(B),1e-5(L) | 12(B),10(L) | 0.2(B),0.3(L) |