Cog-VADU: A Training-Free Cognitive Reasoning Framework for Video Anomaly Detection and Understanding
Organizations: Centre for Vision, Speech and Signal Processing (CVSSP) University of Surrey, UK · Surrey Institute for People-Centred AI (PAI) University of Surrey, UK
Abstract
Video Anomaly Detection (VAD) aims to temporally localize abnormal events in videos. Most existing approaches rely on dataset-specific training and curated annotations, limiting generalization in open-set scenarios. Recent zero-shot methods based on Large Vision- Language Models (LVLMs) alleviate this dependency but often lack temporal continuity and structured reasoning. We propose Cog-VADU, a fully training-free framework that reformulates VAD as a sequential cognitive reasoning task. Cog-VADU introduces Chain-of- Anomaly Detection Thought Prompting (CoADTP), which unrolls an LVLM into a recurrent reasoning chain across video segments. By propagating structured rationales over time, the model maintains implicit temporal memory, enabling robust discrimination between com- plex anomalies and high-motion normal activities. To improve reliability, we further design a cross-modal re-ranking stage that aligns textual rationales with visual embeddings, enforcing semantic consistency and temporal coherence for refined and stable predictions. Extensive experiments on multiple public VAD benchmarks demonstrate that Cog-VADU achieves competitive zero-shot performance. Moreover, cross-model evaluations show that CoADTP consistently enhances reasoning-based anomaly detection in a model-agnostic manner, pro- viding interpretable and generalizable anomaly understanding for real-world applications.
Figures & tables
| Prompt | Prior | Reason | Temporal | AUC | AP | Acc@0.5 |
| Prompt 1 (Minimal) | ✗ | ✗ | ✗ | 0.7111 | 0.2150 | 78.28 |
| Prompt 2 (Format-Structured) | ✓ | ✗ | ✗ | 0.7403 | 0.2119 | 57.93 |
| CoADTP(Raw Stage) | ✓ | ✓ | ✓ | 0.7888 | 0.3773 | 82.41 |
| Method | BLEU | METEOR | ROUGE-L | ||||||
| C | E | V | C | E | V | C | E | V | |
| Training + Instruction-Tuned VAD MLLMs | |||||||||
| Hawk Tang et al. (2024) | 0.320 | 0.165 | 0.202 | 0.228 | 0.191 | 0.196 | 0.156 | 0.104 | 0.114 |
| Holmes-VAD Zhang et al. (2024b) | 0.514 | 0.318 | 0.306 | 0.224 | 0.237 | 0.211 | 0.235 | 0.164 | 0.161 |
| Holmes-VAU Zhang et al. (2025b) | 0.913 | 0.804 | 0.566 | 0.190 | 0.165 | 0.121 | 0.329 | 0.370 | 0.355 |
| Open-source MLLMs | |||||||||
| Method | Judg. | Desc. | Anal. | Capt. | Overall |
| Training + instruction-tuned VAD MLLMs | |||||
| HAWK Tang et al. (2024) | 0.078 | 0.075 | 0.071 | 0.077 | 0.077 |
| Holmes-VAD Zhang et al. (2024b) | 0.167 | 0.236 | 0.235 | 0.263 | 0.245 |
| Holmes-VAU Zhang et al. (2025b) (SOTA, in-domain) | 0.652 | 0.297 | 0.479 | 0.430 | 0.431 |
| Open-source MLLMs + CoADTP (zero-shot) | |||||
| Video-ChatGPT Maaz et al. (2024) + CoADTP | 0.400 | 0.170 | 0.084 | 0.104 | 0.149 |
| Method | SORA | OpenSORA | RG2 | VideoLCM | MS-T2 | Avenue | Ped1 | Ped2 |
| MLLM-based VAD (Training+ instruction Tuned ) | ||||||||
| Holmes-VAU Zhang et al. (2025b) | 2.17 | 34.00 | 24.00 | 29.81 | 25.00 | 6.06 | 3.33 | 5.56 |
| Holmes-VAD Zhang et al. (2024b) | 6.52 | 34.00 | 32.00 | 33.56 | 22.92 | 12.12 | 20.00 | 5.56 |
| HAWK Tang et al. (2024) | 24.64 | 52.00 | 44.00 | 36.54 | 50.00 | 36.36 | 36.67 | 38.89 |
| VAD-R1 Huang et al. (2025) | 41.30 | 78.00 | 56.00 | 63.46 | 60.42 | 75.76 | 60.00 | 63.89 |
| Open-Source MLLMs | ||||||||
Appendix figures & tables15 assets
Supplementary material from the paper’s appendix.
Appendix
| Configuration | Reasoning | Anomaly Database | Temporal Feedback | Refine- ment | AUC | AP |
| Base Prompt | ✗ | ✗ | ✗ | ✗ | 0.7111 | 0.2150 |
| + Reasoning only | ✓ | ✗ | ✗ | ✗ | 0.7548 | 0.2488 |
| + Anomaly DB only | ✗ | ✓ | ✗ | ✗ | 0.7406 | 0.3392 |
| + Temporal FB only | ✗ | ✗ | ✓ | ✗ | 0.7439 | 0.3303 |
| + Reasoning + DB | ✓ | ✓ | ✗ | ✗ | 0.7810 | 0.3238 |
| + R + DB + Temporal FB | ✓ | ✓ | ✓ | ✗ | 0.7888 | 0.3773 |
| Corruption Type | Recovery Rate | Median | Mean Max | Mean Final |
| Wrong category (Normal Violence) | 20/20 (100%) | 1 | 0.509 | 0.139 |
| Inverted (Anomaly Normal) | 20/20 (100%) | 1 | 0.446 | 0.180 |
| Hallucinated entity (Masked + Weapon) | 20/20 (100%) | 1 | 0.512 | 0.142 |
| Method | AP (Orig.) | AP (Ours-Ann.) |
| Weakly Supervised Non-Explainable Methods | ||
| UR-DMU Zhou et al. (2023a) [AAAI-23] | 35.59 | 50.84 |
| Explainable VAD Methods | ||
| VERA Ye et al. (2025) [CVPR-25] | 32.83 | 45.12 |
| Holmes-VAU Zhang et al. (2025b) [CVPR-25] | 36.24 | 56.49 |
| Training-Free Explainable Methods | ||