Video Large Language Models (VideoLLMs) have achieved strong performance on video understanding tasks, yet existing benchmarks evaluate only what models predict, leaving the stability of how they generate largely unexamined. We surface a previously underexplored generation failure of VideoLLMs, defined as output repetition, in which the decoder collapses into self-reinforcing loops of repeated phrases or sentences, and present VideoSTF, a benchmarking framework for systematically measuring, stress-testing, and exploiting this failure mode. VideoSTF formalizes repetition with three complementary n-gram-based metrics, ships a standardized testbed of 10,000 diverse videos, and provides a library of controlled temporal stressors. Across 10 advanced VideoLLMs, VideoSTF reveals four key findings: (i) repetition is pervasive on unperturbed videos and stable across commonly used frame counts, with repetition rates up to 91%; (ii) it spans a severity spectrum from mild redundancy to token-cap loops, and is highly amplified by temporal perturbations; (iii) temporal stressors form a practical black-box attack surface, flipping benign videos into repetitive ones with tens of queries and high attack success rates (up to 98%), and (iv) repetition is not explained by visual redundancy, its amplification tracks local temporal disruption, and only repetition penalties reduce it among common mitigations such as top-k sampling, input filtering, and prompt variation, but increasing the penalty weakens visual grounding. VideoSTF reframes generation stability as a useful and complementary evaluation axis for VideoLLMs and provides the tools to study it. The project page is available at https://videostf.github.io/.
Figures & tables
Figure 1 : Examples of normal and repetitive outputs.
Figure 2 : VideoSTF . (a) Framework Components . VideoSTF comprises three n -gram-based repetition metrics, a standardized testbed of 10,000 videos with diverse durations and content categories, and a library of controlled temporal stressors for applying temporal transformations. (b) Evaluation Protocols . VideoSTF assesses output repetition through three tests. Pervasive Testing reveals widespread repetition across VideoLLMs under different frame sampling settings across all three metrics. Temporal Stress Testing shows that temporal transformations amplify repetition. Adversarial Exploitation demonstrates that, through temporal transformations, videos with normal outputs can be efficiently induced to become repetitive with high success rates and few queries.
Model (LLM)
Frames
RR
RI
IE
Model (LLM)
Frames
RR
RI
IE
LLaVA-Video-7B-Qwen2 (Qwen2)
8
3
0.32
0.87
VideoLLaMA2 (Mistral-7B-Instruct-v0.2)
8
13
0.37
0.84
16
5
0.33
0.86
16
15
0.38
0.84
24
10
0.34
0.86
24
16
0.39
0.83
32
7
0.32
0.86
32
15
0.38
0.83
LLaVA-Video-7B-Qwen2-Video-Only (Qwen2)
8
63
0.47
0.80
ShareGPT4Video (Meta-Llama-3-8B-Instruct)
8
91
0.58
0.75
16
67
0.48
0.80
16
85
0.57
0.76
Table 1 : Repetition rates (%), repetition intensity and information entropy of mainstream VideoLLMs .
Figure 3 : Repetition rates (%) on original videos and under temporal transformations.
Figure 4 : RI and IE distributions under original and temporally transformed inputs across VideoLLMs .
Figure 5 : Attack performance of stressors on videos that originally produce non-repetitive outputs.
Model
Text only
Black
Noise
Static
Genuine frames
1
2
4
8
16
24
32
LLaVA-Video-7B-Qwen2
56
2
0
0
9
2
10
3
7
11
8
LLaVA-Video-7B-Qwen2-Video-Only
85
11
0
30
4
23
50
64
68
64
61
ShareGPT4Video
0
0
0
12
26
57
76
87
85
80
84
Qwen3-VL-8B-Instruct
0
0
0
11
14
8
6
9
12
12
10
Molmo2-8B
0
0
2
27
15
22
21
59
63
60
74
Table 2 : Repetition rates (%) under visual-input controls with fixed weights, prompt, and decoding.
Figure 6 : Feature-level changes induced by temporal transformations.
Figure 7 : Repetition rates of ShareGPT4Video under decoding strategies, TRF, and prompts.
Appendix figures & tables26 assets
Supplementary material from the paper’s appendix.
Appendix
Figure 8 : Repetition results of different VideoLLMs under three metrics across varying n .
Model
Benign
Avg Tok
Ordinary Rep
Avg Tok
Severe Rep
Avg Tok
>1
>2
>3
>4
LLaVA-Video-7B-Qwen2
93
131.54
7
188.29
0
-
7
1
0
0
LLaVA-Video-7B-Qwen2-Video-Only
42
171.5
51
253.83
7
1024
58
27
14
8
ShareGPT4Video
18
206.06
65
356.82
17
1024
82
47
22
11
VideoLLaMA2
85
124.87
13
173.23
2
1024
15
7
3
3
Molmo2-8B
28
326.58
49
403.94
23
1024
72
32
13
9
Appendix
Table 3 : Repetition severity statistics (%) across models at 32 frames. Rep denotes repetition, Avg Tok is the average number of output tokens, and column >k gives the share of outputs whose maximum n -gram count exceeds k .
Figure 9 : Repetition severity statistics across models. (a) Distribution of benign, ordinary, and severe (token-cap) outputs. (b) Number of outputs whose maximum n -gram count exceeds τ .
Quality measure
RR
RI
IE
Information density
-0.76
-0.95
0.92
Output length (words)
0.72
0.84
-0.84
Distinct content words
0.35
0.40
-0.41
CLIPScore
0.00
0.03
-0.04
Appendix
Table 4 : Spearman correlation ( ρ ) between quality measures and the repetition metrics.
Group (maximum n -gram count)
Share of outputs (%)
Words
CLIPScore
Distinct content words
Information density
Benign (1)
72.2
145
0.2163
67.0
46.1
Ordinary (2 to 4)
24.7
230
0.2232
78.8
34.3
Severe (more than 4)
3.1
382
0.2198
64.7
16.9
Appendix
Table 5 : Output quality by repetition severity.
Model
Benign description
Repetitive description
LLaVA-Video-7B-Qwen2
86.6
83.6
LLaVA-Video-7B-Qwen2-Video-Only
84.7
84.9
Appendix
Table 6 : NExT-QA multiple-choice accuracy (%) split by whether the description repeats.
Figure 10 : CLIP-based frame similarity of videos with benign vs. repetitive outputs.
Condition
RR
Change
Original (unperturbed)
8
–
2 frames replaced with frames from a different video
17
+9
Gaussian blur on the same 2 frames
10
+2
Gaussian noise on the same 2 frames
9
+1
JPEG compression (quality 10) on the same 2 frames
8
0
Cutout on the same 2 frames
8
0
Appendix
Table 7 : Repetition rates (%) of LLaVA-Video-7B-Qwen2 at 32 frames under frame-level changes.
Condition
LLaVA-Video-7B-Qwen2
LLaVA-Video-7B-Qwen2-Video-Only
Original, 32 frames
8
61
Add 2, 34 frames
18
55
Uniform re-sample, 34 frames
10
49
Delete 2, 30 frames
10
57
Uniform re-sample, 30 frames
14
61
Appendix
Table 8 : Repetition rates (%) of Add and Delete against count-matched uniform re-sampling.
Trans
Δ Rep
Δ Adj Sim
Δ Glob Sim
Δ Rank
Add 1
0.08
-0.0010
8.1E-05
-0.0036
Add 2
0.07
-0.0021
-3.7E-05
-0.00070
Delete 1
0.08
-0.00016
1.9E-05
-0.0038
Delete 2
0.06
-0.00037
6.1E-05
-0.0092
Replace 1
0.05
-0.00099
-0.00017
-0.0065
Replace 2
0.05
-0.0019
0.00025
-0.045
Appendix
Table 9 : Feature-level changes and their correlation with repetition.
Output
First 50 tokens
Last 50 tokens
Repetitive
19.3
9.6
Benign
19.0
9.5
Appendix
Table 10 : Share of decoder attention on visual tokens (%).
Training corpus
Mean words
Caption RI
Phrase rate
Model
RR (%)
RI
LLaVA-Video-178K
1055
0.677
1.98
LLaVA-Video-7B-Qwen2
5
0.33
LLaVA-Video-7B-Qwen2-Video-Only
67
0.48
ShareGPT4Video mix
288
0.387
0.47
ShareGPT4Video
85
0.57
ShareGPT4V
144
0.370
0.02
VideoLLaMA2
15
0.38
LLaVA-Hound
149
0.322
0.00
LLaVA-NeXT-Video-7B-DPO
8
0.37
LLaVA-Instruct-150K
170
0.317
0.03
VideoLLaMA2
15
0.38
Appendix
Table 11 : Caption repetition in training corpora and the corresponding models.
Source
Name
Mean words
RR at 150 words
RR at 300 words
Corpus
LLaVA-Video-178K
1055
25.7
86.4
ShareGPT4Video mix
288
2.7
12.4
ShareGPT4V
144
11.2
–
Model
LLaVA-Video-7B-Qwen2
–
30.0
–
LLaVA-Video-7B-Qwen2-Video-Only
–
59.4
100.0
Appendix
Table 12 : Length-matched repetition rates (%) of training captions and model outputs.
Setting
RR (%)
Severe (%)
Words
CLIPScore
Distinct content words
T0 (greedy, default)
61
12
209
0.2323
72.3
T0.7 p0.9
60
20
227
0.2295
77.5
T0.7 p0.9 r1.1
50
8
212
0.2284
84.6
T0.7 p0.9 r1.2
36
0
224
0.2261
105.3
T0.7 p0.9 r1.3
10
0
225
0.2196
121.1
Appendix
Table 13 : Repetition and description quality under decoding settings.
Setting
RR
CLIPScore
Caption QA
NExT-QA
MVBench
Multiple choice
Open-ended
T0 (greedy)
61
0.2323
63.9
83.3
16.5
58.9
T0.7 p0.9
60
0.2295
64.1
79.8
14.6
56.0
T0.7 p0.9 r1.1
50
0.2284
62.3
79.8
14.6
56.0
T0.7 p0.9 r1.2
36
0.2261
61.2
79.8
14.6
56.0
T0.7 p0.9 r1.3
10
0.2196
59.5
79.8
14.6
56.0
Appendix
Table 14 : Repetition, description quality, and benchmark accuracy (%) under decoding settings.
Alternative view
Recovery rate
Shuffle
21
Reverse
15
2 × speed re-sample
15
Delete 2 frames
12
Cascade over all four views
41
Appendix
Table 15 : Recovery rate (%) of repetitive outputs under alternative temporal views.
Figure 11 : Repetition generalizes across video types, including long, reasoning, instructional, egocentric, and synthetic videos.
Figure 12 : Examples of repetitive outputs generated by LLaVA-Video-7B-Qwen2 and LLaVA-Video-7B-Qwen2-Video-Only , with repeated phrases or sentences highlighted in different colors.
Figure 13 : Examples of repetitive outputs generated by LLaVA-NeXT-Video-7B and LLaVA-NeXT-Video-7B-DPO , with repeated phrases or sentences highlighted in different colors. We observe highly repetitive responses when the sampled frame number falls within [28,32) .
Figure 14 : Examples of repetitive outputs generated by LLaVA-NeXT-Video-32B-Qwen and VideoLLaMA2 , with repeated phrases or sentences highlighted in different colors.
Figure 15 : Examples of repetitive outputs generated by ShareGPT4Video and InternVL3.5-8B , with repeated phrases or sentences highlighted in different colors.
Figure 16 : Examples of repetitive outputs generated by Qwen3-VL-8B-Instruct and Molmo2-8B , with repeated phrases or sentences highlighted in different colors.
Figure 17 : Example video under temporal frame insertion . We show the sampled frames and corresponding outputs from LLaVA-Video-7B-Qwen2-Video-Only at 16 frames for the original video, as well as the videos after transformations of Add 1 frame and Add 2 frames, with repeated phrases or sentences highlighted in different colors.
Figure 18 : Example video under temporal frame deletion . We show the sampled frames and corresponding outputs from LLaVA-Video-7B-Qwen2-Video-Only at 16 frames for the original video, as well as the videos after transformations of Delete 1 frame and Delete 2 frames, with repeated phrases or sentences highlighted in different colors.
Figure 19 : Example video under temporal frame replacement . We show the sampled frames and corresponding outputs from LLaVA-Video-7B-Qwen2-Video-Only at 16 frames for the original video, as well as the videos after transformations of Replace 1 frame and Replace 2 frames, with repeated phrases or sentences highlighted in different colors.
Figure 20 : Example video under temporal reversal and temporal shuffling . We show the sampled frames and corresponding outputs from LLaVA-Video-7B-Qwen2-Video-Only at 16 frames for the original video, as well as the videos after transformations of Reverse and Shuffle , with repeated phrases or sentences highlighted in different colors.
Video-Language Models (VidLMs) achieve strong benchmark scores, yet these scores often hide whether models use the video at all. We show that VidLM failures follow two pathways: some visual signals are never reliably encoded, while others are encoded but overridden by model priors. We introduce REVEAL, a diagnostic stress-test benchmark for quantifying when and why VidLMs under-use visual evidence. REVEAL contains five controlled probes: camera-motion sensitivity, cross-frame integration, video sycophancy, language-only shortcuts, and temporal expectation bias. Together, they test whether models encode basic video signals, combine evidence across frames, and preserve visual evidence against user assertions, language cues, and learned event expectations. Across 12 VidLMs we find systematic failures along both pathways, with most models falling below chance on the binary and six-way probes that humans solve at 78--100% accuracy. Under assertive prompts, a model's output distribution becomes nearly invariant to whether it is shown a real video or random noise, making visual evidence effectively causally inert. We further carry out mechanistic probes to identify where these failures arise in the model pipeline and why visual evidence is lost. REVEAL provides a scalable, human-verified framework for moving beyond aggregate scores toward structured, reproducible evaluation of multimodal reliability.
Sethuraman T, Savya Khosla, Aditi Tiwari +12
University of Illinois Urbana-Champaign, USA · Adobe Research, USA · Qualtrics, USA +6
Recent video generation models can produce highly realistic videos from natural language instructions, with visual quality approaching cinematic standards. Existing evaluation benchmarks, however, predominantly assess visual quality, aesthetic appeal and physical plausibility, while paying limited attention to text, an essential medium for conveying information in everyday scenes. A generated video may appear visually compelling and feature lifelike subjects, yet still render the text within the scene incorrectly. To address this overlooked dimension, we introduce \textbf{VTR-Bench}, a systematic benchmark for evaluating the \textbf{V}isual \textbf{T}ext \textbf{R}endering capabilities of video generation models. VTR-Bench situates text within concrete application scenarios, such as advertisements and scientific videos, with 300 carefully constructed prompts spanning five scenario categories. We develop an automated evaluation pipeline with human alignments that separately assesses text fidelity through carrier-specific transcription and scene and motion requirements through a prompt-specific chain of query. Beyond evaluation, we introduce a \textbf{Keyframe-Guided Agentic Framework} in which a Director agent coordinates image and video generation with visual evaluation, guiding iterative refinement and candidate selection through visual feedback. Experiments on 11 state-of-the-art models reveal widespread difficulties in accurately rendering scene text, with the best-performing model recording an overall word error rate (WER) of 0.250. We further analyze text rendering failures to characterize the challenges faced by current video generation models. These findings highlight visual text rendering as a key challenge for video generation and demonstrate a practical path toward improvement. Code is available at https://github.com/hardenyu21/VTR-Bench.
Yu Huang, Jungang Li, Zhiyuan Wang +8
Department of Data Science, City University of Hong Kong · The Hong Kong University of Science and Technology (Guangzhou) · Westlake University +3
Recent advances in video large language models (VideoLLMs) call for new evaluation protocols and benchmarks for video-language understanding. Inspired by the needle in a haystack test widely used by LLMs, we introduce a novel task of Needle in a Montage (NeMo), which is designed to assess the temporal understanding capabilities of advanced VideoLLMs. Specifically, the proposed task focuses on two fundamental abilities critical for temporal understanding, i.e., retrieval-style long-context recall and temporal grounding. To generate video question answering data for our task, we develop a scalable automated data generation pipeline that facilitates high-quality data synthesis. Built upon the proposed pipeline, we present NeMoBench, a video-language benchmark centered on our task. Specifically, our full set of NeMoBench features 31,378 automatically generated question-answer (QA) pairs from 13,486 videos with various durations ranging from seconds to hours. Experiments demonstrate that our pipeline can reliably and automatically generate high-quality evaluation data, enabling NeMoBench to be continuously updated with the latest videos. We evaluate 20 state-of-the-art models on our benchmark, providing extensive results and key insights into their capabilities and limitations. Our project page is available at: https://lavi-lab.github.io/NeMoBench.
Zi-Yuan Hu, Shuo Liang, Duo Zheng +10
The Chinese University of Hong Kong · Phoenix TV · Stanford University +2