cs.CVOct 8, 2026

VESSI - VLM-Enhanced Support for Surveillance and Investigations

Authors: Saverio Cavasin, Pietro Tedeschi, Mattia Tamiazzo, Alessandro Brighente, Simone Milani, Mauro Conti

Organizations: Department of Mathematics, University of Padua, Padua, Italy · Innovation & Research Center, CY4GATE S.p.A., part of ELT Group, Rome, Italy · Department of Information Engineering, University of Padua, Padua, Italy · Örebro University, Örebro, Sweden

Abstract

Automated video surveillance analysis has become a critical component of intelligence infrastructures and Law Enforcement agencies. Traditional systems lack the semantic module for comprehensive situational awareness and forensic tasks, limiting their ability to interpret events meaningfully or support post-incident investigations. This slows operational insight and increases the burden on human analysts. Recent advances in Vision-Language Models (VLMs) offer promising pathways to bridge this gap. To address this, we propose VLM-Enhanced Support for Surveillance and Investigations (VESSI), a VLM-based framework designed to enhance automated video surveillance analysis through prompt-driven interrogation of video sequences where salient visual features are converted into textual descriptions. We test our framework with four state-of-the-art models. Since most datasets for this task are unlabeled, we also propose the Composite Model Utility Score (CMUS) to assess VLM performance. Experimental results show that our solution substantially improves the analysis capabilities of human operators and enhances the flexibility of automated surveillance systems. In our evaluation, the most reliable model flagged potentially relevant activity in more than 66% of the videos while reducing review time by more than 85%, offering a practical balance between selectivity and efficiency. The model ordering produced by the reference-free CMUS evaluation was reproduced by the normal-video CMUS evaluation and matched the false-positive-rate ordering obtained from 5,909 manually referenced frames. This agreement supports the operational use of the score within the evaluated setting.

Figures & tables

Appendix figures & tables5 assets

Supplementary material from the paper’s appendix.

Appendix

Explore similar work

CardsList
  1. Can Vision-Language Models Analyze Human-Centered Video? Mapping Model Capabilities and Human-AI Collaborative Workflows

    Sep 23, 2026Xiyuan Shen, Jiuyang Lyu, Seokhyun Hwang +4Video UnderstandingVideo-Language Model Evaluation

  2. Parser-Free VLM Verification for Federated Weakly Supervised Video Anomaly Detection

    Sep 7, 2026Sébastien Thuau, Amira Gran, Siba Haidar +1Multimodal Anomaly DetectionWeakly Supervised Video Anomaly Detection

  3. BLUE: Semantics-Preserving Video Compression for Efficient Vision-Language Surveillance Analytics

    Jul 21, 2026Shubham Baid, Akash James, Sahil Chachra +2Efficient VLM InferenceVideo Compression