Multimodal large language models learn mostly from paired image-text data or annotated video, and raw web video is rarely used to further train an existing language model. We study whether raw video, with no captions and no text loss, can serve as mid-training data for a pretrained language model. Frames are encoded into continuous visual tokens, and the language model learns to predict the next visual token. We mid-train Qwen3-1.7B on raw clips from YT-Temporal-1B and then apply the same image-text instruction tuning to it and to the model without mid-training, so that the two differ only in mid-training. The mid-trained model scores 2.9 points higher on average across four video benchmarks and 5.1 points higher across ten image benchmarks, spanning perception, document, and chart tasks. Text performance is preserved even though mid-training includes no text, with an average of 48.9 across 14 text benchmarks compared with 48.0 for the model without mid-training. Analyses across training show that the image and video gains emerge within 30% of training and plateau thereafter, varying by less than 0.5 points. Predicting captions fails to outperform next-visual-token prediction, demonstrating that video mid-training can remain purely self-supervised without the computational overhead or labeling noise of automated captioning.
Figures & tables
Figure 1: Overview of continuous video mid-training. Ordered video frames x1:T are encoded and projected into continuous token representations. Spatial patch grids are serialized with row-delimiter tokens and concatenated into a unified sequence V . A causal language model processes the continuous visual stream without captions or text supervision, and a prediction head rω regresses the subsequent visual representation vi+1 from the hidden state hi . A stop-gradient operator sg(⋅) is applied to the regression target so that gradients do not flow through it.
Figure 2: Captioning prompt for the caption baselines. Qwen3-VL-30B-A3B receives two consecutive frames and writes one caption per frame, following a prompt adapted from the progress-aware captioning of Xue et al. (2025) . Each caption describes the action in its own frame without referring to the other frame, and each <image> is replaced by the visual tokens of that frame.
Benchmark
No Mid-Training
Video Mid-Training
Text-Rich Image Understanding
ChartQA
43.76
51.96
DocVQA
32.13
38.57
InfographicVQA
22.31
25.28
TextVQA
45.49
51.05
General Perception & Diagrams
Table 1: Effect of video mid-training on video benchmarks. No Mid-Training denotes the language model instruction-tuned without video mid-training, and Video Mid-Training denotes the model mid-trained for one epoch (64.4B visual tokens) and then given the same instruction tuning. All scores are accuracies in percent, and the best value in each row is in bold.
Benchmark
No Mid-Training
Video Mid-Training
NExT-QA
59.04
62.18
VideoMME
42.85
45.15
TempCompass
51.58
53.04
EgoSchema
45.20
50.00
Average
49.67
52.59
Table 2: Effect of video mid-training on image benchmarks. Columns follow Table 1 . All scores are accuracies in percent, with OCRBench divided by ten, and the best value in each row is in bold.
Benchmark
Qwen3-1.7B
No Mid-Training
Video Mid-Training
Commonsense & General Knowledge
OpenBookQA
25.99
26.79
26.59
WinoGrande
60.53
61.87
61.64
HellaSwag
45.18
45.26
45.68
SIQA
44.03
44.18
44.49
PIQA
71.47
72.34
72.17
Table 3: Effect of video mid-training on text benchmarks. Qwen3-1.7B is the language model before any multimodal training. All scores are accuracies in percent, and the best value in each row is in bold.
Figure 3: Progress over mid-training. Each point corresponds to an intermediate mid-training checkpoint followed by identical instruction tuning (0.3, 0.5, and 1.0 epochs), plotted against cumulative visual tokens processed during mid-training. The point at zero represents the baseline instruction-tuned model without video mid-training. Figure 4 reports individual benchmark trajectories.
Figure 4: Performance trajectory across mid-training on each benchmark. Each point corresponds to an intermediate mid-training checkpoint after subsequent instruction tuning, plotted against cumulative visual tokens processed during mid-training. The point at zero represents the baseline model without mid-training. In text benchmark panels, the dashed line indicates the original Qwen3-1.7B before multimodal training.
Objective
Video
Image
Text
No mid-training
49.67
48.72
48.03
Current caption
51.28
48.77
47.44
Next caption
52.51
53.54
47.40
Visual next token
52.59
53.80
48.89
Table 4: Comparison of mid-training objectives on YT-1B. All models are instruction-tuned on LLaVA-OneVision-Data, and No mid-training is the same model instruction-tuned without mid-training. The best value in each column is in bold.
Video-text pretraining has achieved remarkable progress through the scaling of models and datasets, yet the quality of language supervision remains underexplored. Existing web-scale datasets often provide only a single sparse caption per video that fails to capture rich spatiotemporal semantics, while directly using captioning models can generate noisy descriptions. We propose a large-scale multimodal large language model-based supervision generation framework that improves supervision diversity, fidelity, and semantic coverage. Starting from 10 million videos, our approach generates multi-view captions (MVC) through complementary summary and detailed captions, reasoning-based refinement, and semantic positive caption generation. To effectively exploit supervision at different granularities, we further introduce a granularity-aware text representation with separate CLS tokens for summary and detailed views. We pretrain video-text models using the resulting supervision corpus and evaluate them across standard, fine-grained and detailed text-to-video retrieval benchmarks. Our approach consistently improves both zero-shot and fine-tuned performance while using smaller pretraining corpora than existing methods, demonstrating the importance of rich and complementary textual supervision for video-text pretraining. Project page: https://rvandeghen.github.io/mvc/
Fida M. Thoker, Renaud Vandeghen, Karen Sanchez +2
King Abdullah University of Science and Technology (KAUST) · University of Li`ege
Recent advancements in chain-of-thought (CoT) reasoning have shown promise in enhancing video understanding and reasoning capabilities of multimodal large language models (MLLMs). However, existing CoT-based MLLMs require labor-intensive CoT annotations and incur substantial training and inference overhead. While visual latent reasoning has emerged as a more efficient alternative, existing methods primarily focus on image tasks and heavily rely on additional supervision signals for visual latent generation (e.g., CoT traces, auxiliary images, or fine-grained annotations), limiting their scalability and transferability to video tasks. To bridge this gap, we introduce VideoLatent, a novel MLLM equipped with a latent injection module tailored for video understanding and reasoning. Specifically, VideoLatent learns to perform visual latent reasoning using a new latent self-forcing training paradigm, which comprises latent alignment and latent diversity objectives, and relies solely on standard video-question-answer triplets. Extensive experiments across 14 benchmarks demonstrate that our model consistently outperforms existing standard and latent MLLMs on general video understanding and complex video reasoning. Compared with Video-R1, our VideoLatent achieves superior computational efficiency, reducing training/inference overhead by ∼6×/∼68×. Moreover, experiments demonstrate that our method has strong generalizability to different MLLM backbones and different model scales.
Video multimodal large language models (MLLMs) can describe what happens in a video, but rarely identify when the supporting evidence occurs. We study generalist video temporal grounding, in which one model predicts a variable-cardinality set of evidence intervals across video lengths, domains, query forms, and viewpoints. Existing training strategies are misaligned with this set-valued task: long-video labels often rely on brittle one-pass annotation, while reinforcement-learning rewards either fail to distinguish non-overlapping predictions or require fragile segment matching. TimeLens2 treats temporal evidence as an interval set throughout supervision and optimization. TimeLens2-93K constructs reliable multi-span supervision through caption-derived proposals, independent localization, cross-agent consensus, semantic verification, and boundary refinement. Our temporal Wasserstein reward computes exact one-dimensional W1 between uniform distributions over merged interval supports, providing dense, matching-free feedback under unequal cardinalities and equivalent fragmentation; temporal IoU complements it with precise-overlap feedback. Across seven benchmarks, TimeLens2-2B outperforms all size-matched baselines on every benchmark, while the 4B and 8B variants achieve state-of-the-art performance, surpassing open-source models with up to 397B parameters. The 2B, 4B, and 8B variants improve over their Qwen3-VL backbones by 14.2, 13.0, and 18.1 mIoU points, respectively.
Yuhan Zhu, Changlian Ma, Xiangyu Zeng +12
Nanjing University · Shanghai AI Laboratory · Shanghai Jiao Tong University +3