PARSEE-VAD: Efficient Training-Free Online Video Anomaly Detection via Proposition-Aware Reasoning and Streaming Evidence Escalation
Authors: Ji Wang, Shuangqing Zhang, Guo-Sen Xie, Fang Zhao
Organizations: Global Institute of Future Technology Shanghai Jiao Tong University Shanghai, 200240, China · School of Intelligence Science and Technology Nanjing University Suzhou, 215163, China · School of Computer Science and Engineering Nanjing University of Science and Technology Nanjing, 210094, China
Training-free online video anomaly detection (VAD) with frozen multimodal language models faces two coupled challenges: extracting reliable current-window semantics under causal and computational constraints, and maintaining temporal continuity without repeatedly transmitting high-dimensional history. Encoding history through text can compress visual evidence and introduce semantic bias, whereas retaining visual history expands multimodal context. We introduce PARSEE-VAD, a two-module framework that separates semantic evidence acquisition from score-state evolution. Proposition-Aware Reasoning (PAR) extracts structured propositional evidence from the current causal window and conditionally activates more specific queries when coarse evidence warrants further refinement. By sharing a reusable causal visual prefix across queries, PAR reduces redundant computation through selective execution. Streaming Evidence Escalation (SEE) maps the acquired proposition evidence into a compact score-domain event state through current evidence escalation, then propagates only the resulting bounded state across decisions to support temporal continuity. Experiments on four benchmarks demonstrate strong training-free online performance while selective routing reduces specialist computation and score-state propagation remains sparse. These results support a current-first principle for streaming multimodal inference: resolve present semantics first, then use compact historical state only to repair residual continuity gaps.
Figures & tables
Figure 1: Motivation for PARSEE-VAD. Left: causal event-state ambiguity under no-future and online-budget constraints, including transient evidence gaps. Middle: inference-interface bottlenecks from scalar scoring and high-dimensional textual or visual memory. Right: PARSEE-VAD separates proposition-based semantic evidence acquisition from compact score-state evolution.
Figure 2: Overview of PARSEE-VAD. One causal visual-prefix state is reused across proposition queries. PA-readout yields order-balanced margins ℓk(t) , and PA-routing selectively acquires P3/P4. SEE converts the available proposition evidence into the current-evidence state Lt through routed escalation f , then carries bounded score-derived state across decisions to emit st . The stream uses a 60-frame (2-s) decision stride.
Method
Training- free
Online
UCF-Crime
XD-Violence
MSAD
UBnormal
AUC (%)
AUC (%)
AP (%)
AUC (%)
AP (%)
AUC (%)
Sultani et al. (2018)
✗
✗
77.92
–
73.20
–
–
50.30
GODS Wang and Cherian (2019)
✗
✗
70.46
61.56
–
–
–
–
RTFM Tian et al. (2021)
✗
✗
83.31
–
77.81
86.70
66.30
64.94
RTFM (online) † Tian et al. (2021)
✗
✓
80.63
–
72.60
–
–
–
CLIP-TSA Joo et al. (2023)
✗
✗
87.58
–
82.19
–
–
–
Table 1: Comparison with existing video anomaly detection methods. ✓ and ✗ indicate whether each method is training-free or online. Here, online indicates that the reported inference uses only observations available by each decision time. PARSEE-VAD headline metrics use completed-interval benchmark mapping; the corresponding release-time availability view is reported in Table 12 . Provenance of the UBnormal entries is documented in Appendix E.1 .
Method
D (s)
C (s)
C/D
REWARD ( Karim et al., 2024 )
6.4
0.500
0.078
MoniTor ( Yang et al., 2025a )
0.6
5.900
9.83
Flashback ( Lee et al., 2025 )
1.0
0.713
0.713
SphereVAD (online) † ( Huang et al., 2026 )
0.167
0.067
0.40
MACD ‡ ( Ouyang et al., 2026 )
1.0
0.520
0.52
PARSEE-VAD (ours)
2.0
1.372
0.686
Table 2: Published online operating points.
Configuration
PA-routing
Score state
Spec. q/w
UCF AUC
MSAD AUC
MSAD AP
Generated (no PA-readout)
–
–
0.000
80.115
88.211
78.071
Q2 only
–
Q2
0.000
84.872
90.275
79.073
+ Q3 (no specialists)
–
Lt
0.000
84.825
90.273
79.978
+ P3/P4 (non-routed)
×
Lt
2.000
–
90.245
81.155
+ P3/P4 (PA-routed)
✓
Lt
0.565
85.071
90.518
81.970
Full PARSEE-VAD
✓
st
0.565
85.146
90.551
82.024
Table 3: Stage-wise ablation of PARSEE-VAD at 512sq.
Figure 3: Efficiency analysis on MSAD. Left: visual-budget accuracy–compute trade-off. Right: latency decomposition for generated scoring, Q2-only, non-routed all-probe, and routed PARSEE inference.
Figure 4: Mechanistic analysis. (a) Representative suppression and reinforcement cases with routed P3/P4 evidence and PARSEE-VAD scores. (b) Distributions of nonzero score-state adjustments. Active is the fraction of decisions with a nonzero adjustment, n counts active adjustments, and Aligned is the fraction of active adjustments whose sign moves the score in the local ground-truth-consistent direction.
Appendix figures & tables14 assets
Supplementary material from the paper’s appendix.
Appendix
Inference strategy
Model s/window ↓
Total s/window ↓
Model speedup ↑
Mean headwise logit MSE ↓
Independent full forward
2.494
4.006
1.00 ×
0.0000
Native prefix reuse (ours)
0.839
2.383
2.97 ×
0.0214
Appendix
Table 4: Shared-prefix reuse: efficiency, routing fidelity, and score impact on MSAD.
Variant
Specialists / win.
Model mean
Total mean
Total med.
Total P95
(s)
(s)
(s)
(s)
Q2 only
0.000
0.802
1.271
1.190
1.687
Dense all-probe
2.000
1.073
1.581
1.499
1.988
PARSEE routed
0.565
0.884
1.372
1.378
1.741
Appendix
Table 5: Full-set MSAD runtime at 512sq.
Dataset
Operating point
AUC (%) ↑
AP (%) ↑
Mean E2E (s) ↓
P95 E2E (s) ↓
MSAD
9F / 60F stride
90.551
82.024
1.372
1.741
MSAD
4F / 30F stride
90.423
83.059
0.753
0.931
UCF-Crime
9F / 60F stride
85.146
–
–
–
UCF-Crime
4F / 30F stride
84.567
–
0.485
0.658
Appendix
Table 6: Canonical and low-latency PARSEE-VAD operating points. The 4F/30F rows use the same model, 512sq budget, PA-routing, and SEE rules as the main system. Matching full-set UCF-Crime timing for the canonical 9F/60F setting was not retained, so those timing cells are left blank.
(a) Current-escalation strength α
α
UCF AUC
MSAD AUC
MSAD AP
Comment
0.25
85.052
90.557
81.266
weaker local update
0.50
85.132
90.632
82.085
0.625
85.144
90.602
82.101
0.75 (used)
85.146
90.551
82.024
deployed
0.875
85.134
90.457
81.851
Appendix
Table 7: Sensitivity of PAR acquisition and SEE current escalation.
Q2 only
PARSEE-VAD
Δ
Backbone
Size
AUC ↑
AP ↑
AUC ↑
AP ↑
AUC
AP
Qwen3.5-2B
2B
88.499
77.869
87.514
78.881
-0.985
+1.012
VideoLLaMA3-7B
7B
88.986
79.071
89.311
81.162
+0.325
+2.090
Qwen3.5-9B
9B
90.275
79.073
90.551
82.024
+0.276
+2.951
Appendix
Table 8: Backbone robustness on MSAD at 512sq.
Dataset
Current update (%)
Temporal maintenance (%)
Final valley gate (%)
UCF-Crime
23.8
1.1
0.32
MSAD
43.8
2.9
0.69
XD-Violence
40.7
2.9
0.56
Appendix
Table 9: SEE intervention density across datasets.
Configuration
Specialists / win.
Total s / win.
AUC
AP
(a) Compute gating versus evidence selection; final score st
PA-routed acquisition
0.565
1.372
90.551
82.024
Always compute + routed mask
2.000
1.581
90.551
82.024
Always compute + dense evidence
2.000
1.581
90.280
81.233
(b) Budget-matched acquisition; current-evidence state Lt
Uniform random, 1,000 seeds
0.565
–
90.147±0.116
80.312±0.508
Appendix
Table 10: MSAD routing audit on a shared 512sq proposition bank.
Policy
Specialists / win.
AUC
AP
Dense all specialists
2.000
67.537
25.476
Uniform random
0.527
68.213±0.224
28.952±0.654
Generic uncertainty
0.527
68.090
31.488
PAR semantic routing
0.527
68.779
34.418
Appendix
Table 11: Cross-dataset routing diagnostic on UCF-Crime.
UCF AUC
XD-Violence
MSAD
Stage
C / A
AUC C / A
AP C / A
AUC C / A
AP C / A
Q2
84.872 / 83.953
92.640 / 90.774
76.047 / 71.465
90.275 / 88.118
79.073 / 76.257
Lt (current)
85.071 / 84.148
92.940 / 91.041
78.876 / 74.015
90.518 / 88.242
81.970 / 78.697
st (final)
85.146 / 84.208
92.999 / 91.082
79.016 / 74.104
90.551 / 88.236
82.024 / 78.669
Appendix
Table 12: Completed-interval and anchor-aligned score availability.
Dataset
Contrast
Δ AUC (95% CI)
Δ AP (95% CI)
UCF-Crime
Q2 →Lt
+0.199 [ −0.098 , +0.536]
–
Lt→st
+0.075 [+0.031, +0.129]
–
XD-Violence
Q2 →Lt
+0.300 [ −0.099 , +0.738]
+2.828 [+0.734, +5.155]
Lt→st
+0.060 [+0.034, +0.088]
+0.140 [+0.074, +0.210]
MSAD
Q2 →Lt
+0.243 [ −0.505 , +1.446]
+2.897 [ −1.204 , +7.466]
Lt→st
+0.032 [ −0.008 , +0.079]
+0.054 [ −0.046 , +0.168]
Appendix
Table 13: Paired video-level bootstrap of score-state increments.
Q2 only
PARSEE-VAD
Δ
Dataset
AUC
AP
AUC
AP
AUC
AP
UCF-Crime
84.43
34.52
84.69
38.67
+0.25
+4.14
XD-Violence
92.24
74.74
92.73
78.35
+0.49
+3.61
MSAD
89.63
78.92
90.11
82.47
+0.48
+3.54
Appendix
Table 14: Decision-anchor stage decomposition.
Forward only
Reverse only
Balanced
Flip (%)
Dataset
AUC
AP
AUC
AP
AUC
AP
UCF-Crime
84.33
33.28
84.06
34.98
84.43
34.52
10.5
XD-Violence
92.06
73.98
91.85
73.93
92.24
74.74
24.1
MSAD
89.92
78.79
89.01
78.83
89.63
78.92
7.9
Appendix
Table 15: Option-order sensitivity of the Q2 proposition readout.
Current-evidence update rule
UCF-Crime
XD-Violence
MSAD
AUC
AP
AUC
AP
AUC
AP
Signed specialist correction (main)
84.69
38.67
92.73
78.35
90.11
82.47
Positive-only specialist reinforcement
84.54
36.09
93.01
78.46
90.10
82.49
Appendix
Table 16: Specialist evidence semantics at decision anchors.
Temporal rule
UCF AUC
XD AUC
XD AP
MSAD AUC
MSAD AP
Current-evidence state Lt
85.071
92.940
78.876
90.518
81.970
Causal MA3
85.630
92.651
77.606
89.405
79.504
Causal EMA ( β=0.5 )
85.987
93.348
79.620
89.752
80.290
Causal Max-3
85.560
92.221
75.230
89.891
79.325
Final score st
85.146
92.999
79.016
90.551
82.024
Appendix
Table 17: Temporal maintenance and parameter sensitivity.
Video anomaly detection (VAD) aims to localize anomalous events in untrimmed videos. Vision-language models (VLMs) provide rich visual understanding for training-free VAD, but existing approaches impose restrictive interfaces between visual understanding and anomaly scoring. Caption-based pipelines compress visual evidence into text, potentially discarding subtle cues, while direct numerical generation forces the model to express its judgment through a small set of predefined scores. Such interfaces can obscure subtle differences in anomaly severity, causing visually distinct clips to receive similar representations or scores and thereby limiting the resolution of anomaly ranking. We propose \textbf{Probe-VAD}, an ordinal binary-probing framework that directly probes severity preferences from a frozen VLM. Given raw video clips, Probe-VAD queries ten ordered severity thresholds and extracts constrained \textit{YES}/\textit{NO} continuation likelihoods. Their normalized preferences form a cumulative severity profile, from which tail evidence is aggregated into a continuous anomaly score, with isotonic projection enforcing ordinal consistency. Experiments on public VAD benchmarks demonstrate superior performance with low computational cost. Probe-VAD provides a simple interface for translating frozen VLM visual understanding into continuous, rank-sensitive anomaly scores without task-specific training or caption-based compression. Code is available at: https://github.com/yvestine/COVAS-VAD.
Jiawei Gu, Qilin Zhao, Tengkuo Guo +6
School of Intelligence Science and Technology, Nanjing University, Suzhou 215163, China · Faculty of Life Science and Medicine, School of Medicine and Health, Harbin Institute of Technology, Harbin 150001, China · School of Instrumentation Science and Engineering, Harbin Institute of Technology, Harbin 150001, China +2
Existing Video Anomaly Detection (VAD) methods typically rely on task-specific training, leading to strong domain dependency and high training costs. Moreover, most existing methods output only scalar anomaly scores, providing limited insight into why specific events are considered abnormal. Recent advances in Vision-Language Models (VLMs) have enabled both anomaly detection and human-interpretable reasoning. However, many VLM-based approaches still require additional training steps (e.g., instruction tuning or verbalized learning) or external Large Language Models (LLMs), incurring further training costs and inference overhead. To address these challenges, we propose CoReVAD, a contextual reasoning framework for training-free video anomaly detection that operates with a single frozen VLM. CoReVAD directly generates anomaly scores and temporal descriptions from the VLM. To mitigate noise in generative outputs, we introduce a Local Response Cleaning (LRC) module based on local vision-text alignment. Furthermore, global temporal context and progression are incorporated through softmax-based refinement, Gaussian smoothing, and position weighting. Experiments on UCF-Crime and XD-Violence demonstrate that CoReVAD achieves competitive performance among training-free methods while providing reliable and interpretable explanations. Our official code is available at: https://github.com/Muk-00/CoReVAD
Hyeongmuk Lim, Youngbum Hur
Department of Industrial Engineering, Inha University, Incheon, Republic of Korea
Video Anomaly Detection (VAD) aims to temporally localize abnormal events in videos. Most existing approaches rely on dataset-specific training and curated annotations, limiting generalization in open-set scenarios. Recent zero-shot methods based on Large Vision- Language Models (LVLMs) alleviate this dependency but often lack temporal continuity and structured reasoning. We propose Cog-VADU, a fully training-free framework that reformulates VAD as a sequential cognitive reasoning task. Cog-VADU introduces Chain-of- Anomaly Detection Thought Prompting (CoADTP), which unrolls an LVLM into a recurrent reasoning chain across video segments. By propagating structured rationales over time, the model maintains implicit temporal memory, enabling robust discrimination between com- plex anomalies and high-motion normal activities. To improve reliability, we further design a cross-modal re-ranking stage that aligns textual rationales with visual embeddings, enforcing semantic consistency and temporal coherence for refined and stable predictions. Extensive experiments on multiple public VAD benchmarks demonstrate that Cog-VADU achieves competitive zero-shot performance. Moreover, cross-model evaluations show that CoADTP consistently enhances reasoning-based anomaly detection in a model-agnostic manner, pro- viding interpretable and generalizable anomaly understanding for real-world applications.
Mohd Ubaid Wani, Sara Atito, Josef Kittler +1
Centre for Vision, Speech and Signal Processing (CVSSP) University of Surrey, UK · Surrey Institute for People-Centred AI (PAI) University of Surrey, UK