Multimodal large language models learn mostly from paired image-text data or annotated video, and raw web video is rarely used to further train an existing language model. We study whether raw video, with no captions and no text loss, can serve as mid-training data for a pretrained language model. Frames are encoded into continuous visual tokens, and the language model learns to predict the next visual token. We mid-train Qwen3-1.7B on raw clips from YT-Temporal-1B and then apply the same image-text instruction tuning to it and to the model without mid-training, so that the two differ only in mid-training. The mid-trained model scores 2.9 points higher on average across four video benchmarks and 5.1 points higher across ten image benchmarks, spanning perception, document, and chart tasks. Text performance is preserved even though mid-training includes no text, with an average of 48.9 across 14 text benchmarks compared with 48.0 for the model without mid-training. Analyses across training show that the image and video gains emerge within 30% of training and plateau thereafter, varying by less than 0.5 points. Predicting captions fails to outperform next-visual-token prediction, demonstrating that video mid-training can remain purely self-supervised without the computational overhead or labeling noise of automated captioning.
Figures & tables
Figure 1: Overview of continuous video mid-training. Ordered video frames x1:T are encoded and projected into continuous token representations. Spatial patch grids are serialized with row-delimiter tokens and concatenated into a unified sequence V . A causal language model processes the continuous visual stream without captions or text supervision, and a prediction head rω regresses the subsequent visual representation vi+1 from the hidden state hi . A stop-gradient operator sg(⋅) is applied to the regression target so that gradients do not flow through it.
Figure 2: Captioning prompt for the caption baselines. Qwen3-VL-30B-A3B receives two consecutive frames and writes one caption per frame, following a prompt adapted from the progress-aware captioning of Xue et al. (2025) . Each caption describes the action in its own frame without referring to the other frame, and each <image> is replaced by the visual tokens of that frame.
Benchmark
No Mid-Training
Video Mid-Training
Text-Rich Image Understanding
ChartQA
43.76
51.96
DocVQA
32.13
38.57
InfographicVQA
22.31
25.28
TextVQA
45.49
51.05
General Perception & Diagrams
Table 1: Effect of video mid-training on video benchmarks. No Mid-Training denotes the language model instruction-tuned without video mid-training, and Video Mid-Training denotes the model mid-trained for one epoch (64.4B visual tokens) and then given the same instruction tuning. All scores are accuracies in percent, and the best value in each row is in bold.
Benchmark
No Mid-Training
Video Mid-Training
NExT-QA
59.04
62.18
VideoMME
42.85
45.15
TempCompass
51.58
53.04
EgoSchema
45.20
50.00
Average
49.67
52.59
Table 2: Effect of video mid-training on image benchmarks. Columns follow Table 1 . All scores are accuracies in percent, with OCRBench divided by ten, and the best value in each row is in bold.
Benchmark
Qwen3-1.7B
No Mid-Training
Video Mid-Training
Commonsense & General Knowledge
OpenBookQA
25.99
26.79
26.59
WinoGrande
60.53
61.87
61.64
HellaSwag
45.18
45.26
45.68
SIQA
44.03
44.18
44.49
PIQA
71.47
72.34
72.17
Table 3: Effect of video mid-training on text benchmarks. Qwen3-1.7B is the language model before any multimodal training. All scores are accuracies in percent, and the best value in each row is in bold.
Figure 3: Progress over mid-training. Each point corresponds to an intermediate mid-training checkpoint followed by identical instruction tuning (0.3, 0.5, and 1.0 epochs), plotted against cumulative visual tokens processed during mid-training. The point at zero represents the baseline instruction-tuned model without video mid-training. Figure 4 reports individual benchmark trajectories.
Figure 4: Performance trajectory across mid-training on each benchmark. Each point corresponds to an intermediate mid-training checkpoint after subsequent instruction tuning, plotted against cumulative visual tokens processed during mid-training. The point at zero represents the baseline model without mid-training. In text benchmark panels, the dashed line indicates the original Qwen3-1.7B before multimodal training.
Objective
Video
Image
Text
No mid-training
49.67
48.72
48.03
Current caption
51.28
48.77
47.44
Next caption
52.51
53.54
47.40
Visual next token
52.59
53.80
48.89
Table 4: Comparison of mid-training objectives on YT-1B. All models are instruction-tuned on LLaVA-OneVision-Data, and No mid-training is the same model instruction-tuned without mid-training. The best value in each column is in bold.