Multimodal Large Language Models (MLLMs) provide a natural way to make video anomaly detection more explainable. However, their final decisions do not always fully use the discriminative information contained in their hidden states, an issue we refer to as representation--behavior misalignment. We decompose this gap into a capacity component that measures discriminative information never aggregated into the readout position, and a directional component that measures the angular mismatch between the optimal and the native normal--abnormal axis at that position. Across multiple video anomaly detection benchmarks and MLLM backbones the directional component dominates, and residual-stream tracing shows that native-axis separability rises sharply in several mid-to-late attention layers. Because both components are governed by attention rather than MLP updates, we propose Representation--Behavior Alignment (RBA), a parameter-efficient method that adapts those layers using video-level labels alone while updating about 0.012% of the backbone parameters. Experiments on three benchmarks show that RBA improves native-readout performance and better aligns the model's decision direction with discriminative representations, and it produces anomaly decisions and explanations through a single generative process.
Figures & tables
Figure 1: Diagnosis of representation–behavior misalignment on XD-Violence. (a) Native readout versus one-dimensional readouts on frozen hidden states. (b) 2D projection showing directional misalignment between the Fisher-optimal axis wLDA and the native readout axis u .
Figure 2: Layer-wise residual-stream diagnostic. The two Fisher references nearly coincide at every depth and saturate early, whereas native separability increases sharply at the Block 20 attention sublayer. Additive decomposition isolates positive gains from attention and negative interference from MLP sublayers.
Figure 3: Overview of the RBA framework. Probe-derived pseudo-labels supervise low-rank adapters injected into the transition attention layers, optimized via a hinge-margin loss along the fixed native readout axis u .
Method
Venue
Interpretable
UCF AUC
UB AUC
XD AUC
XD AP
Specialized Detectors (Non-interpretable)
VadCLIP [ 28 ]
AAAI’24
✗
88.02
—
—
84.51
π -VAD [ 19 ]
CVPR’25
✗
90.33
—
—
85.37
RefineVAD [ 12 ]
AAAI’26
✗
88.92
—
—
88.66
LAVIDA [ 8 ]
CVPR’26
✗
82.18
76.45
—
90.62
HeadHunt-VAD [ 6 ]
AAAI’26
✗
87.03
—
—
82.63
Table 1: Quantitative comparison on VAD benchmarks (%). ✓ and ✗ denote native support for natural language explanations.
Backbone / Dataset
Pfull
Plast
Pprobe
Pread
Δcap
Δdir
Dir. Ratio
Aligned
cos(wLDA,u)
Qwen3.5-9B / XD
0.9646
0.9588
0.8756
0.7286
0.0058
0.2302
97.5%
0.8302
0.238 → 0.774
Qwen3.5-9B / UCF
0.9168
0.9120
0.8847
0.7714
0.0048
0.1406
96.7%
0.8701
0.287 → 0.789
Qwen3.5-9B / UB
0.8974
0.8910
0.8523
0.7290
0.0064
0.1620
96.2%
0.7684
0.263 → 0.746
InternVL3.5-8B / XD
0.9251
0.9180
0.8734
0.6493
0.0071
0.2687
97.4%
0.7842
0.164 → 0.728
Gemma-4-12B / XD
0.9386
0.9320
0.8938
0.6980
0.0066
0.2340
97.3%
0.8058
0.201 → 0.756
Table 2: Two-factor gap decomposition and alignment gains across benchmarks and backbones, reporting Frame-AP for XD and Frame-AUC for UCF/UB. Pfull and Plast are retrospective Fisher ceilings used only for decomposition, whereas Pprobe is the weakly supervised attainable ceiling against which the recovered fraction of Δdir is reported in the text. The final column gives the cosine between the Fisher-optimal axis and the native readout axis before and after adaptation.
Table 6
Figure 4: Qualitative and geometric analysis. (a)(b) Temporal anomaly scores and corresponding explanations before and after RBA. (c)(d) Final-layer feature distributions along the native readout axis u , reflecting the downstream effect of adapting Blocks 20–21.
Appendix figures & tables18 assets
Supplementary material from the paper’s appendix.
Appendix
Figure 5: Pseudo-label layer diagnosis on the XD-Violence training set (Qwen3.5-9B). (a) Native Frame-AP against pseudo-labels across depth. (b) Pseudo-label discriminability dread′ and cumulative sublayer updates, showing a sharp phase transition at the Block 20 attention sublayer.
Figure 6: Pseudo-label layer diagnosis on the UCF-Crime training set (Qwen3.5-9B). (a) Native Frame-AUC against pseudo-labels across depth. (b) Pseudo-label discriminability dread′ , consistently locating the transition at Block 20.
Figure 7: Pseudo-label layer diagnosis on the UBnormal training set (Qwen3.5-9B). (a) Native Frame-AUC against pseudo-labels across depth. (b) Pseudo-label discriminability dread′ , verifying the identical Block 20 attention transition under synthetic anomaly scenarios.
Pooling strategy
Pseudo-P
Pseudo-R
XD AP
XD AUC
UCF AUC
UB AUC
Mean Pooling
52.41
89.67
75.14
91.23
81.38
69.82
Max Pooling
84.16
41.29
78.43
92.47
83.94
72.18
Random- K Pooling
56.83
74.38
76.07
91.62
82.16
70.49
Top- K MIL
81.79
79.34
83.02
94.30
87.01
76.84
Dual-Memory Units
83.27
81.08
83.46
94.51
87.23
77.12
Appendix
Table 5: Ablation of MIL pooling strategies on downstream RBA performance (%). Pseudo-label precision and recall are diagnostic quantities computed against training-set temporal annotations; these annotations are not used for optimization.
Placement
Sublayer
Params
Frame-AP
Native Readout
–
0
0.7286
Blocks 5–6
Post-Attention
≈0.012%
0.7314
Blocks 17–18
Post-Attention
≈0.012%
0.7976
Blocks 20–21
Post-Attention
≈0.012%
0.8302
Blocks 30–31
Post-Attention
≈0.012%
0.7738
Blocks 20–21
Post-MLP
≈0.024%
0.7415
Appendix
Table 6: Adapter placement ablation on XD-Violence. All attention placements use approximately the same trainable parameter budget.
Rank r
Params
XD AP
XD AUC
UCF AUC
UB AUC
1
≈0.001%
77.83
92.14
82.47
73.19
4
≈0.003%
80.46
93.18
84.82
74.93
8
≈0.006%
82.17
93.84
86.29
76.15
16
≈0.012%
83.02
94.30
87.01
76.84
32
≈0.024%
83.09
94.36
87.04
76.81
64
≈0.048%
82.93
94.25
86.97
76.68
Appendix
Table 7: LoRA rank ablation in Blocks 20–21 with α/r=2.0 .
Target projections
Params
XD AP
UCF AUC
UB AUC
Query, key, and value ( q,k,v )
≈0.006%
80.67
85.14
75.32
Attention output ( o )
≈0.002%
79.83
84.67
74.89
Standard attention ( q,k,v,o )
≈0.008%
82.14
86.38
76.23
All attention projections
≈0.012%
83.02
87.01
76.84
Appendix
Table 8: Ablation of target attention projections in Blocks 20–21 using r=16 and α=32 .
Margin M
0.5
1.0
1.5
2.0
2.5
3.0
Frame-AP
78.94
81.23
82.68
83.02
82.89
81.95
Frame-AUC
92.81
93.74
94.18
94.30
94.22
93.80
Appendix
Table 9: Sensitivity of the hinge margin M on XD-Violence.
Training objective
Frame-AP
Frame-AUC
cos(wprobe′,u)
Zero-Shot Base Readout
72.86
90.37
0.238
LBCE (hard labels)
79.82
93.12
0.684
LKL (soft distribution)
78.45
92.68
0.641
LMSE (logit regression)
77.19
91.85
0.612
Lalign+Lexplain
72.79
90.41
0.298
Lalign
83.02
94.30
0.781
Appendix
Table 10: Comparison of alignment objectives on XD-Violence. The metric cos(wprobe′,u) evaluates directional alignment between the adapted probe direction and the fixed native axis.
Figure 8: Layer-wise residual-stream diagnostic on Qwen3.5-9B using alternative word-level readout tokens ( Yes / No ) on XD-Violence. Left: three-tier readout trajectories across network depth. Right: native discriminability dread′ and cumulative sublayer contributions.
Figure 9: Layer-wise residual-stream diagnostic on Qwen3.5-9B using Safe / Risky readout tokens on XD-Violence. Left: three-tier readout trajectories across network depth. Right: native discriminability dread′ and cumulative sublayer contributions.
Figure 10: Layer-wise residual-stream diagnostic on InternVL3.5-8B on XD-Violence using the main-experiment readout tokens. Left: three-tier readout trajectories across network depth. Right: native discriminability dread′ and cumulative sublayer contributions.
Figure 11: Layer-wise residual-stream diagnostic on Gemma 4 on XD-Violence using the main-experiment readout tokens. Left: three-tier readout trajectories across network depth. Right: native discriminability dread′ and cumulative sublayer contributions.
Method
Judge
Consist.
F+
R−
Quality
Zero-Shot Qwen
GPT-5.4 mini
78.2%
94.5%
86.1%
3.63
Zero-Shot Qwen
Gemini 3.6 Flash
79.1%
93.8%
85.4%
3.57
Probe + Caption
GPT-5.4 mini
43.2%
71.3%
94.4%
2.85
Probe + Caption
Gemini 3.6 Flash
44.0%
70.8%
93.9%
2.78
RBA
GPT-5.4 mini
91.4%
96.3%
36.7%
4.12
RBA
Gemini 3.6 Flash
92.0%
95.9%
35.8%
4.09
Appendix
Table 11: Decision–explanation evaluation on fixed diagnostic sets. Consistency and F+ are higher-is-better, while R− is lower-is-better.
Dataset
Metric
Base Readout
RBA
Absolute Gain
XD-Violence
Frame-AP
72.86
83.02 ± 0.17
+10.16 ± 0.17
Frame-AUC
90.37
94.30 ± 0.11
+3.93 ± 0.11
UCF-Crime
Frame-AUC
77.14
87.01 ± 0.19
+9.87 ± 0.19
UBnormal
Frame-AUC
72.90
76.84 ± 0.22
+3.94 ± 0.22
Appendix
Table 12: Multi-seed evaluation over five independent runs. Metrics are reported in percent. The base readout is deterministic under the fixed prompt and greedy decoding protocol.
Figure 12: Quantitative and qualitative comparison for Case Group I. Left: temporal anomaly-score trajectories produced by the zero-shot Qwen3.5 readout and RBA, together with frame-level ground truth and the decision threshold. Right: explanations generated at representative frames. RBA suppresses a false alarm in the normal traffic video ( 0.62→0.05 ) and recovers a missed anomaly in the abnormal crowd video ( 0.09→0.96 ).
Figure 13: Quantitative and qualitative comparison for Case Group II. RBA corrects a false positive in a normal sports video ( 0.82→0.04 ) and a false negative in an abnormal dense-crowd video ( 0.03→0.96 ). The generated explanations are shown alongside the temporal score trajectories.
Group
Scene
GT
Frame
ZS Score
RBA Score
ZS Margin
RBA Margin
Δm
I
Nighttime traffic
Normal
3711
0.62
0.05
−0.12
+0.45
+0.57
I
Outdoor crowd
Abnormal
2136
0.09
0.96
−0.41
+0.46
+0.87
II
Sports activity
Normal
2919
0.82
0.04
−0.32
+0.46
+0.78
II
Dense crowd
Abnormal
6360
0.03
0.96
−0.47
+0.46
+0.93
Mean
−0.33
+0.46
+0.79
Appendix
Table 13: Quantitative summary of the representative cases. ZS denotes the zero-shot native readout. A positive signed margin indicates a correct prediction under the 0.5 threshold.
Vision-language models (VLMs) have recently emerged as a promising paradigm for video anomaly detection (VAD) due to their strong visual reasoning ability and natural language-based explainability. In this paper, we aim to address a key limitation of such pipelines, which perform segment-level inference independently owing to token constraints and reason without structured temporal context, allowing VLMs to interpret anomalies as deviations from evolving video dynamics rather than producing fragmented predictions and explanations. To specify, we propose a context-aware framework named LATERN, which reformulates VAD as a temporal evidence aggregation process. LATERN consists of two complementary modules: Context-Aware Anomaly Scoring (CEA) and Recursive Evidence Aggregation (REA). CEA introduces a novel image-grounded memory mechanism, which selectively chooses historical content via frame diversity and visual-textual alignment as expanded context to help generate reliable anomaly scores. Building upon these scores, REA performs recursive temporal aggregation to identify coherent anomaly intervals and produce event-level decisions and explanations grounded in visual-textual evidence. Extensive experiments on challenging benchmarks, including UCF-Crime and XD-Violence, show that LATERN enhances detection accuracy and explanation consistency for frozen VLMs during test time, while generating temporally coherent and semantically grounded event-level explanations.
Recent video anomaly detection research has expanded rapidly with an emphasis on general models of normality intended to work across many different scenes. While this focus has led to improvements in scalability and multi-scene generalization, it has also shifted the field away from modeling the scene-specific and context-dependent nature of normal behavior. Contemporary approaches frequently rely on video-level weak supervision and opaque pretrained representations from multi-modal large language models (MLLMs), which encourage models to respond to familiar semantic anomaly categories rather than to deviations from the normal patterns of a particular environment. This trend suppresses spatial localization, introduces semantic bias, and reduces anomaly detection to a form of action recognition. In this paper, we examine whether these prevailing formulations align with the core requirements of real-world VAD, which is typically performed within a single scene where normality is determined by local geometry, semantics, and activity patterns. Through targeted visual analyses and empirical evaluations, we demonstrate the practical consequences of these limitations and show that meaningful progress in VAD requires renewed focus on single-scene, spatially-aware, and explainable formulations that capture the nuanced structure of normality within individual environments.
Furkan Mumcu, Michael J. Jones, Anoop Cherian +1
University of South Florida · Mitsubishi Electric Research Laboratories (MERL)
Video anomaly detection (VAD) with multimodal large language models has shown strong potential, yet most existing methods still depend on large-scale annotations or expert-designed priors, limiting their ability to acquire anomaly knowledge with as little human intervention as possible. To address this, we propose Linguistic Relative Policy Optimization (LRPO), which distills group-relative semantic advantages from multiple reasoning trajectories into a linguistically expressed anomaly experience prior, and adapts the model by injecting this prior into the context to steer its output distribution without any parameter updates. LRPO builds two complementary experience representations: general experience captures transferable anomaly preferences across scenarios, while scenario experience models context-dependent anomaly rules for targeted refinement. To further improve the learned experience, we introduce an anomaly alignment reward that guides trajectory optimization to match human risk preferences and reinforce temporally grounded reasoning. Extensive experiments on XD-Violence, UCF-Crime, and UBnormal demonstrate that LRPO significantly outperforms existing state-of-the-art methods under tuning-free settings.
Jiaxu Leng, Jiankang Zheng, Mengjingcheng Mo +4
School of Computer Science and Technology, Chongqing University of Posts and Telecommunications, Chongqing, China · Chongqing College of Artificial Intelligence, Chongqing, China