VideoSTF: Stress-Testing Output Repetition in Video Large Language Models
Organizations: National University of Singapore, Singapore · University of New South Wales, Australia · CSIRO’s Data61, Australia
Abstract
Video Large Language Models (VideoLLMs) have achieved strong performance on video understanding tasks, yet existing benchmarks evaluate only what models predict, leaving the stability of how they generate largely unexamined. We surface a previously underexplored generation failure of VideoLLMs, defined as output repetition, in which the decoder collapses into self-reinforcing loops of repeated phrases or sentences, and present VideoSTF, a benchmarking framework for systematically measuring, stress-testing, and exploiting this failure mode. VideoSTF formalizes repetition with three complementary -gram-based metrics, ships a standardized testbed of 10,000 diverse videos, and provides a library of controlled temporal stressors. Across 10 advanced VideoLLMs, VideoSTF reveals four key findings: (i) repetition is pervasive on unperturbed videos and stable across commonly used frame counts, with repetition rates up to 91%; (ii) it spans a severity spectrum from mild redundancy to token-cap loops, and is highly amplified by temporal perturbations; (iii) temporal stressors form a practical black-box attack surface, flipping benign videos into repetitive ones with tens of queries and high attack success rates (up to 98%), and (iv) repetition is not explained by visual redundancy, its amplification tracks local temporal disruption, and only repetition penalties reduce it among common mitigations such as top- sampling, input filtering, and prompt variation, but increasing the penalty weakens visual grounding. VideoSTF reframes generation stability as a useful and complementary evaluation axis for VideoLLMs and provides the tools to study it. The project page is available at https://videostf.github.io/.
Figures & tables
| Model (LLM) | Frames | RR | RI | IE | Model (LLM) | Frames | RR | RI | IE |
| LLaVA-Video-7B-Qwen2 (Qwen2) | 8 | 3 | 0.32 | 0.87 | VideoLLaMA2 (Mistral-7B-Instruct-v0.2) | 8 | 13 | 0.37 | 0.84 |
| 16 | 5 | 0.33 | 0.86 | 16 | 15 | 0.38 | 0.84 | ||
| 24 | 10 | 0.34 | 0.86 | 24 | 16 | 0.39 | 0.83 | ||
| 32 | 7 | 0.32 | 0.86 | 32 | 15 | 0.38 | 0.83 | ||
| LLaVA-Video-7B-Qwen2-Video-Only (Qwen2) | 8 | 63 | 0.47 | 0.80 | ShareGPT4Video (Meta-Llama-3-8B-Instruct) | 8 | 91 | 0.58 | 0.75 |
| 16 | 67 | 0.48 | 0.80 | 16 | 85 | 0.57 | 0.76 |
| Model | Text only | Black | Noise | Static | Genuine frames | ||||||
| 1 | 2 | 4 | 8 | 16 | 24 | 32 | |||||
| LLaVA-Video-7B-Qwen2 | 56 | 2 | 0 | 0 | 9 | 2 | 10 | 3 | 7 | 11 | 8 |
| LLaVA-Video-7B-Qwen2-Video-Only | 85 | 11 | 0 | 30 | 4 | 23 | 50 | 64 | 68 | 64 | 61 |
| ShareGPT4Video | 0 | 0 | 0 | 12 | 26 | 57 | 76 | 87 | 85 | 80 | 84 |
| Qwen3-VL-8B-Instruct | 0 | 0 | 0 | 11 | 14 | 8 | 6 | 9 | 12 | 12 | 10 |
| Molmo2-8B | 0 | 0 | 2 | 27 | 15 | 22 | 21 | 59 | 63 | 60 | 74 |
Appendix figures & tables26 assets
Supplementary material from the paper’s appendix.
Appendix
| Model | Benign | Avg Tok | Ordinary Rep | Avg Tok | Severe Rep | Avg Tok | >1 | >2 | >3 | >4 |
| LLaVA-Video-7B-Qwen2 | 93 | 131.54 | 7 | 188.29 | 0 | - | 7 | 1 | 0 | 0 |
| LLaVA-Video-7B-Qwen2-Video-Only | 42 | 171.5 | 51 | 253.83 | 7 | 1024 | 58 | 27 | 14 | 8 |
| ShareGPT4Video | 18 | 206.06 | 65 | 356.82 | 17 | 1024 | 82 | 47 | 22 | 11 |
| VideoLLaMA2 | 85 | 124.87 | 13 | 173.23 | 2 | 1024 | 15 | 7 | 3 | 3 |
| Molmo2-8B | 28 | 326.58 | 49 | 403.94 | 23 | 1024 | 72 | 32 | 13 | 9 |
| Quality measure | RR | RI | IE |
| Information density | -0.76 | -0.95 | 0.92 |
| Output length (words) | 0.72 | 0.84 | -0.84 |
| Distinct content words | 0.35 | 0.40 | -0.41 |
| CLIPScore | 0.00 | 0.03 | -0.04 |
| Group (maximum -gram count) | Share of outputs (%) | Words | CLIPScore | Distinct content words | Information density |
| Benign (1) | 72.2 | 145 | 0.2163 | 67.0 | 46.1 |
| Ordinary (2 to 4) | 24.7 | 230 | 0.2232 | 78.8 | 34.3 |
| Severe (more than 4) | 3.1 | 382 | 0.2198 | 64.7 | 16.9 |
| Model | Benign description | Repetitive description |
| LLaVA-Video-7B-Qwen2 | 86.6 | 83.6 |
| LLaVA-Video-7B-Qwen2-Video-Only | 84.7 | 84.9 |
| Condition | RR | Change |
| Original (unperturbed) | 8 | – |
| 2 frames replaced with frames from a different video | 17 | +9 |
| Gaussian blur on the same 2 frames | 10 | +2 |
| Gaussian noise on the same 2 frames | 9 | +1 |
| JPEG compression (quality 10) on the same 2 frames | 8 | 0 |
| Cutout on the same 2 frames | 8 | 0 |
| Condition | LLaVA-Video-7B-Qwen2 | LLaVA-Video-7B-Qwen2-Video-Only |
| Original, 32 frames | 8 | 61 |
| Add 2, 34 frames | 18 | 55 |
| Uniform re-sample, 34 frames | 10 | 49 |
| Delete 2, 30 frames | 10 | 57 |
| Uniform re-sample, 30 frames | 14 | 61 |
| Trans | Rep | Adj Sim | Glob Sim | Rank |
| Add 1 | 0.08 | -0.0010 | 8.1E-05 | -0.0036 |
| Add 2 | 0.07 | -0.0021 | -3.7E-05 | -0.00070 |
| Delete 1 | 0.08 | -0.00016 | 1.9E-05 | -0.0038 |
| Delete 2 | 0.06 | -0.00037 | 6.1E-05 | -0.0092 |
| Replace 1 | 0.05 | -0.00099 | -0.00017 | -0.0065 |
| Replace 2 | 0.05 | -0.0019 | 0.00025 | -0.045 |
| Output | First 50 tokens | Last 50 tokens |
| Repetitive | 19.3 | 9.6 |
| Benign | 19.0 | 9.5 |
| Training corpus | Mean words | Caption RI | Phrase rate | Model | RR (%) | RI |
| LLaVA-Video-178K | 1055 | 0.677 | 1.98 | LLaVA-Video-7B-Qwen2 | 5 | 0.33 |
| LLaVA-Video-7B-Qwen2-Video-Only | 67 | 0.48 | ||||
| ShareGPT4Video mix | 288 | 0.387 | 0.47 | ShareGPT4Video | 85 | 0.57 |
| ShareGPT4V | 144 | 0.370 | 0.02 | VideoLLaMA2 | 15 | 0.38 |
| LLaVA-Hound | 149 | 0.322 | 0.00 | LLaVA-NeXT-Video-7B-DPO | 8 | 0.37 |
| LLaVA-Instruct-150K | 170 | 0.317 | 0.03 | VideoLLaMA2 | 15 | 0.38 |
| Source | Name | Mean words | RR at 150 words | RR at 300 words |
| Corpus | LLaVA-Video-178K | 1055 | 25.7 | 86.4 |
| ShareGPT4Video mix | 288 | 2.7 | 12.4 | |
| ShareGPT4V | 144 | 11.2 | – | |
| Model | LLaVA-Video-7B-Qwen2 | – | 30.0 | – |
| LLaVA-Video-7B-Qwen2-Video-Only | – | 59.4 | 100.0 |
| Setting | RR (%) | Severe (%) | Words | CLIPScore | Distinct content words |
| T0 (greedy, default) | 61 | 12 | 209 | 0.2323 | 72.3 |
| T0.7 p0.9 | 60 | 20 | 227 | 0.2295 | 77.5 |
| T0.7 p0.9 r1.1 | 50 | 8 | 212 | 0.2284 | 84.6 |
| T0.7 p0.9 r1.2 | 36 | 0 | 224 | 0.2261 | 105.3 |
| T0.7 p0.9 r1.3 | 10 | 0 | 225 | 0.2196 | 121.1 |
| Setting | RR | CLIPScore | Caption QA | NExT-QA | MVBench | |
| Multiple choice | Open-ended | |||||
| T0 (greedy) | 61 | 0.2323 | 63.9 | 83.3 | 16.5 | 58.9 |
| T0.7 p0.9 | 60 | 0.2295 | 64.1 | 79.8 | 14.6 | 56.0 |
| T0.7 p0.9 r1.1 | 50 | 0.2284 | 62.3 | 79.8 | 14.6 | 56.0 |
| T0.7 p0.9 r1.2 | 36 | 0.2261 | 61.2 | 79.8 | 14.6 | 56.0 |
| T0.7 p0.9 r1.3 | 10 | 0.2196 | 59.5 | 79.8 | 14.6 | 56.0 |
| Alternative view | Recovery rate |
| Shuffle | 21 |
| Reverse | 15 |
| 2 speed re-sample | 15 |
| Delete 2 frames | 12 |
| Cascade over all four views | 41 |