cs.CVOct 8, 2026

A Unified Score Matching Paradigm for Video Anomaly Detection and Anticipation

Authors: Congqi Cao, Zhenhe Liang, Hanwen Zhang, Yifan Zhao, Qinyi Lv, Lingtong Min, Yanning Zhang

Organizations: National Engineering Laboratory for Integrated Aero-Space-Ground-Ocean Big Data Application Technology, School of Computer Science, Northwestern Polytechnical University, Xi’an 710129, China · School of Electronics and Information, Northwestern Polytechnical University, Xi’an 710129, China

Abstract

Video anomaly detection (VAD) is a fundamental and safety-critical task in computer vision. Recent generative approaches detect anomalies from a distributional perspective, but remain limited by local anomaly modes. Meanwhile, video anomaly anticipation (VAA), as a proactive extension beyond post-hoc detection, introduces additional challenges. In particular, the contrastive inference paradigm in VAD, which relies on ground-truth frames, is not applicable to VAA, hindering its development. To address these challenges, we propose a unified score-driven framework, termed Uni-DSM, based on denoising score matching (DSM), which models anomaly patterns through likelihood estimation and score functions over the learned data distribution. Within this unified framework, we adopt a shared noise-conditioned score transformer backbone with scene-dependent embeddings and motion-aware weighting for distribution-level modeling. Instead of introducing separate architectures, Uni-DSM unifies VAD and VAA through different inference and supervision paradigms built upon the same score-based formulation. For VAD, we instantiate an autoregressive denoising score matching (ADSM) mechanism, which progressively accumulates anomalous evidence via autoregressive denoising, enabling enhanced perception of local modes beyond visual cues. For VAA, we extend the same architecture by incorporating a lightweight auxiliary decoder and a novel self-distilled denoising score matching (SDSM) mechanism. By constructing supervision from output discrepancies instead of relying on unavailable future ground truth, our method achieves efficient training suitable or early anomaly anticipation. Extensive experiments on multiple benchmark datasets demonstrate state-of-the-art performance in both VAD and VAA while maintaining high efficiency, establishing a unified and scalable pipeline from anomaly detection to anticipation.

Figures & tables

Explore similar work

CardsList
  1. Glance, Scrutinize, and Think: Advancing Video Anomaly Detection from Training-Free to Agentic Reasoning

    Aug 7, 2026Shibo Gao, Peipei Yang, Xu-Yao Zhang +1Anomaly LocalizationVideo Anomaly Detection

  2. Probe-VAD: Ordinal Likelihood Probing for Training-Free Video Anomaly Detection

    Sep 15, 2026Jiawei Gu, Qilin Zhao, Tengkuo Guo +6Video Anomaly DetectionZero-Shot Anomaly Detection

  3. Cog-VADU: A Training-Free Cognitive Reasoning Framework for Video Anomaly Detection and Understanding

    Oct 1, 2026Mohd Ubaid Wani, Sara Atito, Josef Kittler +1CoT ReasoningInterpretable Anomaly Detection