cs.CVOct 8, 2026

VESSI - VLM-Enhanced Support for Surveillance and Investigations

Authors: Saverio Cavasin, Pietro Tedeschi, Mattia Tamiazzo, Alessandro Brighente, Simone Milani, Mauro Conti

Organizations: Department of Mathematics, University of Padua, Padua, Italy · Innovation & Research Center, CY4GATE S.p.A., part of ELT Group, Rome, Italy · Department of Information Engineering, University of Padua, Padua, Italy · Örebro University, Örebro, Sweden

Abstract

Automated video surveillance analysis has become a critical component of intelligence infrastructures and Law Enforcement agencies. Traditional systems lack the semantic module for comprehensive situational awareness and forensic tasks, limiting their ability to interpret events meaningfully or support post-incident investigations. This slows operational insight and increases the burden on human analysts. Recent advances in Vision-Language Models (VLMs) offer promising pathways to bridge this gap. To address this, we propose VLM-Enhanced Support for Surveillance and Investigations (VESSI), a VLM-based framework designed to enhance automated video surveillance analysis through prompt-driven interrogation of video sequences where salient visual features are converted into textual descriptions. We test our framework with four state-of-the-art models. Since most datasets for this task are unlabeled, we also propose the Composite Model Utility Score (CMUS) to assess VLM performance. Experimental results show that our solution substantially improves the analysis capabilities of human operators and enhances the flexibility of automated surveillance systems. In our evaluation, the most reliable model flagged potentially relevant activity in more than 66% of the videos while reducing review time by more than 85%, offering a practical balance between selectivity and efficiency. The model ordering produced by the reference-free CMUS evaluation was reproduced by the normal-video CMUS evaluation and matched the false-positive-rate ordering obtained from 5,909 manually referenced frames. This agreement supports the operational use of the score within the evaluated setting.

Figures & tables

Appendix figures & tables5 assets

Supplementary material from the paper’s appendix.

Appendix

Explore similar work

Sep 23, 2026cs.CV

Can Vision-Language Models Analyze Human-Centered Video? Mapping Model Capabilities and Human-AI Collaborative Workflows

Video provides a rich record of human behavior, interaction, and situated contexts, offering important evidence for understanding people and conducting human-centered research. As vision-language models (VLMs) become increasingly capable of analyzing video, they offer opportunities to automate this traditionally human-intensive process. Yet a central question remains: when can VLMs analyze human-centered video independently, and when does reliable analysis still require human involvement? To address this question, we first characterize video analysis practices in human-centered research. We systematically analyze all 1,702 CHI 2026 full papers and identify 125 that annotate videos. Through iterative coding, we derive a five-dimensional taxonomy spanning analytic purpose, viewpoint, phenomenon, reasoning requirement, and annotation authority. Grounded in recurring annotation tasks captured by this taxonomy, we construct a benchmark of 15 representative tasks from open datasets to map the capabilities and limitations of a general-purpose VLM. We examine the division of labor between humans and VLMs by comparing three annotation workflows: VLM alone, human alone, and human verification of VLM outputs. Across tasks, VLM-alone annotation approaches human accuracy on average (HNS = 97.0, where 100 denotes human-alone performance), demonstrating substantial potential to automate human-centered video analysis. Human verification achieves the highest accuracy (HNS = 121.5) while reducing human annotation time by 48.9% and monetary cost by 31.3%-44.5% relative to human-alone annotation. Our findings connect real-world human-centered video analysis tasks and current VLM capabilities, and clarify how human-AI collaboration can make VLM-assisted analysis reliable and efficient.
Sep 7, 2026cs.CV

Parser-Free VLM Verification for Federated Weakly Supervised Video Anomaly Detection

How can vision-language models help video anomaly detection (VAD) when surveillance data remain distributed, weakly labeled, and resource-constrained? Most weakly supervised VAD methods assume centralized training; recent VLM-based extensions further rely on dense inference, generated explanations, or additional adaptation. We introduce a lightweight federated MIL-VLM cascade in which only a compact MIL scorer is trained across clients, while a frozen VLM verifies high-scoring suspect segments post hoc. We study two VLM feedback interfaces: parsed text-generation decisions and a logit-based interface that extracts a continuous anomaly score from next-token Yes/No probabilities. Experiments on UCF-Crime with InternVL3.5-2B and Qwen3-VL-2B-Instruct show that text-generation verification can improve frame-level AUC after diagnostic temporal post-processing, but remains sensitive to prompts, parsers, model choice, and smoothing. In contrast, the logit interface provides a fixed parser-free signal that improves both frame-level AUC and frame-level AP over the MIL baseline across both VLMs, without temporal post-processing in its main configuration. Since suspect segments are updated independently once available, next-token logit feedback provides a simple segment-local alternative to text-generation verification.
Jul 21, 2026eess.IV

BLUE: Semantics-Preserving Video Compression for Efficient Vision-Language Surveillance Analytics

Continuous surveillance video creates a growing storage, transmission, and inference burden for enterprise video analytics systems. While modern codecs such as H.265 reduce bitrate for human-viewable video, aggressive compression can degrade downstream computer-vision performance and does not necessarily reduce the number of vision-language model (VLM) inference calls required for semantic video understanding. This paper evaluates BLUE, a fixed-camera surveillance compression approach that suppresses static-background redundancy while preserving foreground activity, for its effect on VLM-based event and anomaly understanding. We compare raw H.265 and BLUE-compressed H.265 video on two surveillance datasets: VIRAT, comprising 227 paired event samples from 106 clips, and CHAD, comprising 54 human-activity anomaly clips. For each pair, the same frame index is evaluated using a VLM captioning pipeline, and outputs are scored against annotation-derived ground truth using a blind judging protocol. The results show no measurable degradation in semantic inference quality. On VIRAT, the mean VLM score remains effectively unchanged between raw H.265 and BLUE, with a mean difference of approximately -0.01 on a 0-10 scale. On CHAD, raw H.265 and BLUE obtain near-equivalent mean scores of 4.31 and 4.26, respectively. Compression saving is also uncorrelated with VLM score change on VIRAT (r = 0.004), indicating that higher BLUE compression does not predict semantic quality loss. Beyond storage reduction, BLUE increases the share of skip-heavy P-frames on CHAD from 1.4% to 53.2%, enabling an estimated 53% reduction in VLM calls through packet-size-based frame skipping. These findings suggest that BLUE functions as a machine-centric compression layer for surveillance video, reducing bandwidth and inference cost while preserving VLM semantic performance.