cs.AISep 28, 2026

Decision Readouts for Text-Mediated Video Anomaly Detection: An Exploratory Evaluation of Jev and Qwen

Authors: Xukui Qin, Youting Wang, Xinjie He, Ziyang Luo, Runxiong Wu, Yan-Syuan Chen, Zhongyao Chu

Organizations: Independent researchers

Abstract

How much does the decision readout matter when video-derived textual evidence is held fixed? We evaluate Jev typed decisions and three Qwen readouts on a sparse development sample of 40 videos and 400 target anchors from UCF-Crime and XD-Violence, each presented as a summary and ordered captions. Each dataset contributes 20 source groups and 200 anchors, including only 10 and 37 positives, respectively. The original five-backend pilot requested 4,000 predictions; Jev Choice returned 776 valid responses out of 800 under the study's strict numerical policy, blocking its full-coverage quality comparison. On XD captions, Jev Noul achieved 75.99% average precision versus 48.47% for Qwen generated probability and 57.81% for the stronger local ordinal-likelihood expectation. The latter paired difference was 18.18 percentage points (95% source-group bootstrap interval 5.53-31.50). UCF did not show a corresponding advantage: caption ROC-AUC was 52.26% for Noul and 65.95% for ordinal likelihood. Both probability readouts had higher, hence worse, UCF Brier scores than the evaluation-prevalence reference of 0.0475. We additionally audit historical LAVAD scores at exactly matched anchors and distinguish response structure from numerical consistency. A binary-likelihood control is missing. These exploratory offline results characterize ranking, probability quality and interface failures; they establish neither a causal typed-interface benefit nor general superiority, calibration or end-to-end acceleration.

Figures & tables

Appendix figures & tables7 assets

Supplementary material from the paper’s appendix.

Appendix

Explore similar work

CardsList
  1. VTAgent: Agentic Keyframe Anchoring for Evidence-Aware Video TextVQA

    May 6, 2026Haibin He, Maoyuan Ye, Jing Zhang +2Long Video Question AnsweringVideo Foundation Models

  2. VIBE: Annotation-Free Video-to-Text Information Bottleneck Evaluation for TL;DR

    May 23, 2025Shenghui Chen, Po-han Li, Sandeep Chinchali +1SummarizationBottlenecks

  3. JEV-as-a-Judge: Accept When Confident, Escalate When Unsure

    Sep 22, 2026Yubo Li, Yidi Miao, Ramayya Krishnan +1Large Language Model JudgesJudges