cs.CVSep 27, 2026

PARSEE-VAD: Efficient Training-Free Online Video Anomaly Detection via Proposition-Aware Reasoning and Streaming Evidence Escalation

Authors: Ji Wang, Shuangqing Zhang, Guo-Sen Xie, Fang Zhao

Organizations: Global Institute of Future Technology Shanghai Jiao Tong University Shanghai, 200240, China · School of Intelligence Science and Technology Nanjing University Suzhou, 215163, China · School of Computer Science and Engineering Nanjing University of Science and Technology Nanjing, 210094, China

Abstract

Training-free online video anomaly detection (VAD) with frozen multimodal language models faces two coupled challenges: extracting reliable current-window semantics under causal and computational constraints, and maintaining temporal continuity without repeatedly transmitting high-dimensional history. Encoding history through text can compress visual evidence and introduce semantic bias, whereas retaining visual history expands multimodal context. We introduce PARSEE-VAD, a two-module framework that separates semantic evidence acquisition from score-state evolution. Proposition-Aware Reasoning (PAR) extracts structured propositional evidence from the current causal window and conditionally activates more specific queries when coarse evidence warrants further refinement. By sharing a reusable causal visual prefix across queries, PAR reduces redundant computation through selective execution. Streaming Evidence Escalation (SEE) maps the acquired proposition evidence into a compact score-domain event state through current evidence escalation, then propagates only the resulting bounded state across decisions to support temporal continuity. Experiments on four benchmarks demonstrate strong training-free online performance while selective routing reduces specialist computation and score-state propagation remains sparse. These results support a current-first principle for streaming multimodal inference: resolve present semantics first, then use compact historical state only to repair residual continuity gaps.

Figures & tables

Appendix figures & tables14 assets

Supplementary material from the paper’s appendix.

Appendix

Explore similar work

CardsList
  1. Probe-VAD: Ordinal Likelihood Probing for Training-Free Video Anomaly Detection

    Sep 15, 2026Jiawei Gu, Qilin Zhao, Tengkuo Guo +6Training-Free Video Anomaly DetectionAnomaly Score

  2. CoReVAD: A Contextual Reasoning Framework for Training-Free Video Anomaly Detection

    May 22, 2026Hyeongmuk Lim, Youngbum HurTraining-Free Video Anomaly DetectionFine-Grained Video Understanding

  3. Cog-VADU: A Training-Free Cognitive Reasoning Framework for Video Anomaly Detection and Understanding

    Oct 1, 2026Mohd Ubaid Wani, Sara Atito, Josef Kittler +1Training-Free Video Anomaly DetectionFine-Grained Video Understanding