Multimodal Large Language Models (MLLMs) provide a natural way to make video anomaly detection more explainable. However, their final decisions do not always fully use the discriminative information contained in their hidden states, an issue we refer to as representation--behavior misalignment. We decompose this gap into a capacity component that measures discriminative information never aggregated into the readout position, and a directional component that measures the angular mismatch between the optimal and the native normal--abnormal axis at that position. Across multiple video anomaly detection benchmarks and MLLM backbones the directional component dominates, and residual-stream tracing shows that native-axis separability rises sharply in several mid-to-late attention layers. Because both components are governed by attention rather than MLP updates, we propose Representation--Behavior Alignment (RBA), a parameter-efficient method that adapts those layers using video-level labels alone while updating about 0.012% of the backbone parameters. Experiments on three benchmarks show that RBA improves native-readout performance and better aligns the model's decision direction with discriminative representations, and it produces anomaly decisions and explanations through a single generative process.
Figures & tables
Figure 1: Diagnosis of representation–behavior misalignment on XD-Violence. (a) Native readout versus one-dimensional readouts on frozen hidden states. (b) 2D projection showing directional misalignment between the Fisher-optimal axis wLDA and the native readout axis u .
Figure 2: Layer-wise residual-stream diagnostic. The two Fisher references nearly coincide at every depth and saturate early, whereas native separability increases sharply at the Block 20 attention sublayer. Additive decomposition isolates positive gains from attention and negative interference from MLP sublayers.
Figure 3: Overview of the RBA framework. Probe-derived pseudo-labels supervise low-rank adapters injected into the transition attention layers, optimized via a hinge-margin loss along the fixed native readout axis u .
Method
Venue
Interpretable
UCF AUC
UB AUC
XD AUC
XD AP
Specialized Detectors (Non-interpretable)
VadCLIP [ 28 ]
AAAI’24
✗
88.02
—
—
84.51
π -VAD [ 19 ]
CVPR’25
✗
90.33
—
—
85.37
RefineVAD [ 12 ]
AAAI’26
✗
88.92
—
—
88.66
LAVIDA [ 8 ]
CVPR’26
✗
82.18
76.45
—
90.62
HeadHunt-VAD [ 6 ]
AAAI’26
✗
87.03
—
—
82.63
Table 1: Quantitative comparison on VAD benchmarks (%). ✓ and ✗ denote native support for natural language explanations.
Backbone / Dataset
Pfull
Plast
Pprobe
Pread
Δcap
Δdir
Dir. Ratio
Aligned
cos(wLDA,u)
Qwen3.5-9B / XD
0.9646
0.9588
0.8756
0.7286
0.0058
0.2302
97.5%
0.8302
0.238 → 0.774
Qwen3.5-9B / UCF
0.9168
0.9120
0.8847
0.7714
0.0048
0.1406
96.7%
0.8701
0.287 → 0.789
Qwen3.5-9B / UB
0.8974
0.8910
0.8523
0.7290
0.0064
0.1620
96.2%
0.7684
0.263 → 0.746
InternVL3.5-8B / XD
0.9251
0.9180
0.8734
0.6493
0.0071
0.2687
97.4%
0.7842
0.164 → 0.728
Gemma-4-12B / XD
0.9386
0.9320
0.8938
0.6980
0.0066
0.2340
97.3%
0.8058
0.201 → 0.756
Table 2: Two-factor gap decomposition and alignment gains across benchmarks and backbones, reporting Frame-AP for XD and Frame-AUC for UCF/UB. Pfull and Plast are retrospective Fisher ceilings used only for decomposition, whereas Pprobe is the weakly supervised attainable ceiling against which the recovered fraction of Δdir is reported in the text. The final column gives the cosine between the Fisher-optimal axis and the native readout axis before and after adaptation.
Table 6
Figure 4: Qualitative and geometric analysis. (a)(b) Temporal anomaly scores and corresponding explanations before and after RBA. (c)(d) Final-layer feature distributions along the native readout axis u , reflecting the downstream effect of adapting Blocks 20–21.
Appendix figures & tables18 assets
Supplementary material from the paper’s appendix.
Appendix
Figure 5: Pseudo-label layer diagnosis on the XD-Violence training set (Qwen3.5-9B). (a) Native Frame-AP against pseudo-labels across depth. (b) Pseudo-label discriminability dread′ and cumulative sublayer updates, showing a sharp phase transition at the Block 20 attention sublayer.
Figure 6: Pseudo-label layer diagnosis on the UCF-Crime training set (Qwen3.5-9B). (a) Native Frame-AUC against pseudo-labels across depth. (b) Pseudo-label discriminability dread′ , consistently locating the transition at Block 20.
Figure 7: Pseudo-label layer diagnosis on the UBnormal training set (Qwen3.5-9B). (a) Native Frame-AUC against pseudo-labels across depth. (b) Pseudo-label discriminability dread′ , verifying the identical Block 20 attention transition under synthetic anomaly scenarios.
Pooling strategy
Pseudo-P
Pseudo-R
XD AP
XD AUC
UCF AUC
UB AUC
Mean Pooling
52.41
89.67
75.14
91.23
81.38
69.82
Max Pooling
84.16
41.29
78.43
92.47
83.94
72.18
Random- K Pooling
56.83
74.38
76.07
91.62
82.16
70.49
Top- K MIL
81.79
79.34
83.02
94.30
87.01
76.84
Dual-Memory Units
83.27
81.08
83.46
94.51
87.23
77.12
Appendix
Table 5: Ablation of MIL pooling strategies on downstream RBA performance (%). Pseudo-label precision and recall are diagnostic quantities computed against training-set temporal annotations; these annotations are not used for optimization.
Placement
Sublayer
Params
Frame-AP
Native Readout
–
0
0.7286
Blocks 5–6
Post-Attention
≈0.012%
0.7314
Blocks 17–18
Post-Attention
≈0.012%
0.7976
Blocks 20–21
Post-Attention
≈0.012%
0.8302
Blocks 30–31
Post-Attention
≈0.012%
0.7738
Blocks 20–21
Post-MLP
≈0.024%
0.7415
Appendix
Table 6: Adapter placement ablation on XD-Violence. All attention placements use approximately the same trainable parameter budget.
Rank r
Params
XD AP
XD AUC
UCF AUC
UB AUC
1
≈0.001%
77.83
92.14
82.47
73.19
4
≈0.003%
80.46
93.18
84.82
74.93
8
≈0.006%
82.17
93.84
86.29
76.15
16
≈0.012%
83.02
94.30
87.01
76.84
32
≈0.024%
83.09
94.36
87.04
76.81
64
≈0.048%
82.93
94.25
86.97
76.68
Appendix
Table 7: LoRA rank ablation in Blocks 20–21 with α/r=2.0 .
Target projections
Params
XD AP
UCF AUC
UB AUC
Query, key, and value ( q,k,v )
≈0.006%
80.67
85.14
75.32
Attention output ( o )
≈0.002%
79.83
84.67
74.89
Standard attention ( q,k,v,o )
≈0.008%
82.14
86.38
76.23
All attention projections
≈0.012%
83.02
87.01
76.84
Appendix
Table 8: Ablation of target attention projections in Blocks 20–21 using r=16 and α=32 .
Margin M
0.5
1.0
1.5
2.0
2.5
3.0
Frame-AP
78.94
81.23
82.68
83.02
82.89
81.95
Frame-AUC
92.81
93.74
94.18
94.30
94.22
93.80
Appendix
Table 9: Sensitivity of the hinge margin M on XD-Violence.
Training objective
Frame-AP
Frame-AUC
cos(wprobe′,u)
Zero-Shot Base Readout
72.86
90.37
0.238
LBCE (hard labels)
79.82
93.12
0.684
LKL (soft distribution)
78.45
92.68
0.641
LMSE (logit regression)
77.19
91.85
0.612
Lalign+Lexplain
72.79
90.41
0.298
Lalign
83.02
94.30
0.781
Appendix
Table 10: Comparison of alignment objectives on XD-Violence. The metric cos(wprobe′,u) evaluates directional alignment between the adapted probe direction and the fixed native axis.
Figure 8: Layer-wise residual-stream diagnostic on Qwen3.5-9B using alternative word-level readout tokens ( Yes / No ) on XD-Violence. Left: three-tier readout trajectories across network depth. Right: native discriminability dread′ and cumulative sublayer contributions.
Figure 9: Layer-wise residual-stream diagnostic on Qwen3.5-9B using Safe / Risky readout tokens on XD-Violence. Left: three-tier readout trajectories across network depth. Right: native discriminability dread′ and cumulative sublayer contributions.
Figure 10: Layer-wise residual-stream diagnostic on InternVL3.5-8B on XD-Violence using the main-experiment readout tokens. Left: three-tier readout trajectories across network depth. Right: native discriminability dread′ and cumulative sublayer contributions.
Figure 11: Layer-wise residual-stream diagnostic on Gemma 4 on XD-Violence using the main-experiment readout tokens. Left: three-tier readout trajectories across network depth. Right: native discriminability dread′ and cumulative sublayer contributions.
Method
Judge
Consist.
F+
R−
Quality
Zero-Shot Qwen
GPT-5.4 mini
78.2%
94.5%
86.1%
3.63
Zero-Shot Qwen
Gemini 3.6 Flash
79.1%
93.8%
85.4%
3.57
Probe + Caption
GPT-5.4 mini
43.2%
71.3%
94.4%
2.85
Probe + Caption
Gemini 3.6 Flash
44.0%
70.8%
93.9%
2.78
RBA
GPT-5.4 mini
91.4%
96.3%
36.7%
4.12
RBA
Gemini 3.6 Flash
92.0%
95.9%
35.8%
4.09
Appendix
Table 11: Decision–explanation evaluation on fixed diagnostic sets. Consistency and F+ are higher-is-better, while R− is lower-is-better.
Dataset
Metric
Base Readout
RBA
Absolute Gain
XD-Violence
Frame-AP
72.86
83.02 ± 0.17
+10.16 ± 0.17
Frame-AUC
90.37
94.30 ± 0.11
+3.93 ± 0.11
UCF-Crime
Frame-AUC
77.14
87.01 ± 0.19
+9.87 ± 0.19
UBnormal
Frame-AUC
72.90
76.84 ± 0.22
+3.94 ± 0.22
Appendix
Table 12: Multi-seed evaluation over five independent runs. Metrics are reported in percent. The base readout is deterministic under the fixed prompt and greedy decoding protocol.
Figure 12: Quantitative and qualitative comparison for Case Group I. Left: temporal anomaly-score trajectories produced by the zero-shot Qwen3.5 readout and RBA, together with frame-level ground truth and the decision threshold. Right: explanations generated at representative frames. RBA suppresses a false alarm in the normal traffic video ( 0.62→0.05 ) and recovers a missed anomaly in the abnormal crowd video ( 0.09→0.96 ).
Figure 13: Quantitative and qualitative comparison for Case Group II. RBA corrects a false positive in a normal sports video ( 0.82→0.04 ) and a false negative in an abnormal dense-crowd video ( 0.03→0.96 ). The generated explanations are shown alongside the temporal score trajectories.
Group
Scene
GT
Frame
ZS Score
RBA Score
ZS Margin
RBA Margin
Δm
I
Nighttime traffic
Normal
3711
0.62
0.05
−0.12
+0.45
+0.57
I
Outdoor crowd
Abnormal
2136
0.09
0.96
−0.41
+0.46
+0.87
II
Sports activity
Normal
2919
0.82
0.04
−0.32
+0.46
+0.78
II
Dense crowd
Abnormal
6360
0.03
0.96
−0.47
+0.46
+0.93
Mean
−0.33
+0.46
+0.79
Appendix
Table 13: Quantitative summary of the representative cases. ZS denotes the zero-shot native readout. A positive signed margin indicates a correct prediction under the 0.5 threshold.
Jul 1, 2026·Jiaxu Leng, Jiankang Zheng, Mengjingcheng Mo +4Complementarity
School of Computer Science and Technology, Chongqing University of Posts and Telecommunications, Chongqing, China · Chongqing College of Artificial Intelligence, Chongqing, China