PARSEE-VAD: Efficient Training-Free Online Video Anomaly Detection via Proposition-Aware Reasoning and Streaming Evidence Escalation
Organizations: Global Institute of Future Technology Shanghai Jiao Tong University Shanghai, 200240, China · School of Intelligence Science and Technology Nanjing University Suzhou, 215163, China · School of Computer Science and Engineering Nanjing University of Science and Technology Nanjing, 210094, China
Abstract
Training-free online video anomaly detection (VAD) with frozen multimodal language models faces two coupled challenges: extracting reliable current-window semantics under causal and computational constraints, and maintaining temporal continuity without repeatedly transmitting high-dimensional history. Encoding history through text can compress visual evidence and introduce semantic bias, whereas retaining visual history expands multimodal context. We introduce PARSEE-VAD, a two-module framework that separates semantic evidence acquisition from score-state evolution. Proposition-Aware Reasoning (PAR) extracts structured propositional evidence from the current causal window and conditionally activates more specific queries when coarse evidence warrants further refinement. By sharing a reusable causal visual prefix across queries, PAR reduces redundant computation through selective execution. Streaming Evidence Escalation (SEE) maps the acquired proposition evidence into a compact score-domain event state through current evidence escalation, then propagates only the resulting bounded state across decisions to support temporal continuity. Experiments on four benchmarks demonstrate strong training-free online performance while selective routing reduces specialist computation and score-state propagation remains sparse. These results support a current-first principle for streaming multimodal inference: resolve present semantics first, then use compact historical state only to repair residual continuity gaps.
Figures & tables
| Method | Training- free | Online | UCF-Crime | XD-Violence | MSAD | UBnormal | ||
| AUC (%) | AUC (%) | AP (%) | AUC (%) | AP (%) | AUC (%) | |||
| Sultani et al. (2018) | ✗ | ✗ | 77.92 | – | 73.20 | – | – | 50.30 |
| GODS Wang and Cherian (2019) | ✗ | ✗ | 70.46 | 61.56 | – | – | – | – |
| RTFM Tian et al. (2021) | ✗ | ✗ | 83.31 | – | 77.81 | 86.70 | 66.30 | 64.94 |
| RTFM (online) † Tian et al. (2021) | ✗ | ✓ | 80.63 | – | 72.60 | – | – | – |
| CLIP-TSA Joo et al. (2023) | ✗ | ✗ | 87.58 | – | 82.19 | – | – | – |
| Method | (s) | (s) | |
|---|---|---|---|
| REWARD ( Karim et al., 2024 ) | 6.4 | 0.500 | 0.078 |
| MoniTor ( Yang et al., 2025a ) | 0.6 | 5.900 | 9.83 |
| Flashback ( Lee et al., 2025 ) | 1.0 | 0.713 | 0.713 |
| SphereVAD (online) † ( Huang et al., 2026 ) | 0.167 | 0.067 | 0.40 |
| MACD ‡ ( Ouyang et al., 2026 ) | 1.0 | 0.520 | 0.52 |
| PARSEE-VAD (ours) | 2.0 | 1.372 | 0.686 |
| Configuration | PA-routing | Score state | Spec. q/w | UCF AUC | MSAD AUC | MSAD AP |
|---|---|---|---|---|---|---|
| Generated (no PA-readout) | – | – | 0.000 | 80.115 | 88.211 | 78.071 |
| Q2 only | – | Q2 | 0.000 | 84.872 | 90.275 | 79.073 |
| + Q3 (no specialists) | – | 0.000 | 84.825 | 90.273 | 79.978 | |
| + P3/P4 (non-routed) | 2.000 | – | 90.245 | 81.155 | ||
| + P3/P4 (PA-routed) | ✓ | 0.565 | 85.071 | 90.518 | 81.970 | |
| Full PARSEE-VAD | ✓ | 0.565 | 85.146 | 90.551 | 82.024 |
Appendix figures & tables14 assets
Supplementary material from the paper’s appendix.
Appendix
| Inference strategy | Model s/window | Total s/window | Model speedup | Mean headwise logit MSE |
|---|---|---|---|---|
| Independent full forward | 2.494 | 4.006 | 1.00 | 0.0000 |
| Native prefix reuse (ours) | 0.839 | 2.383 | 2.97 | 0.0214 |
| Variant | Specialists / win. | Model mean | Total mean | Total med. | Total P95 |
|---|---|---|---|---|---|
| (s) | (s) | (s) | (s) | ||
| Q2 only | 0.000 | 0.802 | 1.271 | 1.190 | 1.687 |
| Dense all-probe | 2.000 | 1.073 | 1.581 | 1.499 | 1.988 |
| PARSEE routed | 0.565 | 0.884 | 1.372 | 1.378 | 1.741 |
| Dataset | Operating point | AUC (%) | AP (%) | Mean E2E (s) | P95 E2E (s) |
|---|---|---|---|---|---|
| MSAD | 9F / 60F stride | 90.551 | 82.024 | 1.372 | 1.741 |
| MSAD | 4F / 30F stride | 90.423 | 83.059 | 0.753 | 0.931 |
| UCF-Crime | 9F / 60F stride | 85.146 | – | – | – |
| UCF-Crime | 4F / 30F stride | 84.567 | – | 0.485 | 0.658 |
| (a) Current-escalation strength | ||||
|---|---|---|---|---|
| UCF AUC | MSAD AUC | MSAD AP | Comment | |
| 0.25 | 85.052 | 90.557 | 81.266 | weaker local update |
| 0.50 | 85.132 | 90.632 | 82.085 | |
| 0.625 | 85.144 | 90.602 | 82.101 | |
| 0.75 (used) | 85.146 | 90.551 | 82.024 | deployed |
| 0.875 | 85.134 | 90.457 | 81.851 | |
| Q2 only | PARSEE-VAD | ||||||
|---|---|---|---|---|---|---|---|
| Backbone | Size | AUC | AP | AUC | AP | AUC | AP |
| Qwen3.5-2B | 2B | 88.499 | 77.869 | 87.514 | 78.881 | -0.985 | +1.012 |
| VideoLLaMA3-7B | 7B | 88.986 | 79.071 | 89.311 | 81.162 | +0.325 | +2.090 |
| Qwen3.5-9B | 9B | 90.275 | 79.073 | 90.551 | 82.024 | +0.276 | +2.951 |
| Dataset | Current update (%) | Temporal maintenance (%) | Final valley gate (%) |
|---|---|---|---|
| UCF-Crime | 23.8 | 1.1 | 0.32 |
| MSAD | 43.8 | 2.9 | 0.69 |
| XD-Violence | 40.7 | 2.9 | 0.56 |
| Configuration | Specialists / win. | Total s / win. | AUC | AP |
|---|---|---|---|---|
| (a) Compute gating versus evidence selection; final score | ||||
| PA-routed acquisition | 0.565 | 1.372 | 90.551 | 82.024 |
| Always compute + routed mask | 2.000 | 1.581 | 90.551 | 82.024 |
| Always compute + dense evidence | 2.000 | 1.581 | 90.280 | 81.233 |
| (b) Budget-matched acquisition; current-evidence state | ||||
| Uniform random, 1,000 seeds | 0.565 | – | ||
| Policy | Specialists / win. | AUC | AP |
|---|---|---|---|
| Dense all specialists | 2.000 | 67.537 | 25.476 |
| Uniform random | 0.527 | ||
| Generic uncertainty | 0.527 | 68.090 | 31.488 |
| PAR semantic routing | 0.527 | 68.779 | 34.418 |
| UCF AUC | XD-Violence | MSAD | |||
|---|---|---|---|---|---|
| Stage | C / A | AUC C / A | AP C / A | AUC C / A | AP C / A |
| Q2 | 84.872 / 83.953 | 92.640 / 90.774 | 76.047 / 71.465 | 90.275 / 88.118 | 79.073 / 76.257 |
| (current) | 85.071 / 84.148 | 92.940 / 91.041 | 78.876 / 74.015 | 90.518 / 88.242 | 81.970 / 78.697 |
| (final) | 85.146 / 84.208 | 92.999 / 91.082 | 79.016 / 74.104 | 90.551 / 88.236 | 82.024 / 78.669 |
| Dataset | Contrast | AUC (95% CI) | AP (95% CI) |
|---|---|---|---|
| UCF-Crime | Q2 | +0.199 [ , +0.536] | – |
| +0.075 [+0.031, +0.129] | – | ||
| XD-Violence | Q2 | +0.300 [ , +0.738] | +2.828 [+0.734, +5.155] |
| +0.060 [+0.034, +0.088] | +0.140 [+0.074, +0.210] | ||
| MSAD | Q2 | +0.243 [ , +1.446] | +2.897 [ , +7.466] |
| +0.032 [ , +0.079] | +0.054 [ , +0.168] |
| Q2 only | PARSEE-VAD | |||||
|---|---|---|---|---|---|---|
| Dataset | AUC | AP | AUC | AP | AUC | AP |
| UCF-Crime | 84.43 | 34.52 | 84.69 | 38.67 | +0.25 | +4.14 |
| XD-Violence | 92.24 | 74.74 | 92.73 | 78.35 | +0.49 | +3.61 |
| MSAD | 89.63 | 78.92 | 90.11 | 82.47 | +0.48 | +3.54 |
| Forward only | Reverse only | Balanced | Flip (%) | ||||
|---|---|---|---|---|---|---|---|
| Dataset | AUC | AP | AUC | AP | AUC | AP | |
| UCF-Crime | 84.33 | 33.28 | 84.06 | 34.98 | 84.43 | 34.52 | 10.5 |
| XD-Violence | 92.06 | 73.98 | 91.85 | 73.93 | 92.24 | 74.74 | 24.1 |
| MSAD | 89.92 | 78.79 | 89.01 | 78.83 | 89.63 | 78.92 | 7.9 |
| Current-evidence update rule | UCF-Crime | XD-Violence | MSAD | |||
|---|---|---|---|---|---|---|
| AUC | AP | AUC | AP | AUC | AP | |
| Signed specialist correction (main) | 84.69 | 38.67 | 92.73 | 78.35 | 90.11 | 82.47 |
| Positive-only specialist reinforcement | 84.54 | 36.09 | 93.01 | 78.46 | 90.10 | 82.49 |
| Temporal rule | UCF AUC | XD AUC | XD AP | MSAD AUC | MSAD AP |
|---|---|---|---|---|---|
| Current-evidence state | 85.071 | 92.940 | 78.876 | 90.518 | 81.970 |
| Causal MA3 | 85.630 | 92.651 | 77.606 | 89.405 | 79.504 |
| Causal EMA ( ) | 85.987 | 93.348 | 79.620 | 89.752 | 80.290 |
| Causal Max-3 | 85.560 | 92.221 | 75.230 | 89.891 | 79.325 |
| Final score | 85.146 | 92.999 | 79.016 | 90.551 | 82.024 |