Intermediate-layer features from multimodal large language models have shown strong potential for video anomaly detection (VAD), yet the origin of their discriminative power remains unclear. We study this question using sparse mixture-of-experts (MoE) models, whose explicit expert structure and sparse activation make their internal computation easier to inspect. With a fully frozen backbone and no additional training, we find that anomaly-related evidence is concentrated in a small set of experts. These experts recur across layers, spontaneously specialize in different anomaly types, and together form a dynamic routing subnetwork. We further show that the output channels most strongly influenced by these experts are also the hidden dimensions that contain the most anomaly-relevant information. Routing statistics can therefore serve as an internal anomaly cue that complements semantic features.Based on these findings, we propose RoMod, an efficient VAD framework trained with only 5% of weakly labeled videos. RoMod includes a Routing-Modulated Fusion module, RoMF, and a Routing-aware Temporal Network, RoTN. RoMF uses routing signals to adaptively recalibrate hidden semantic channels. Its design also prevents the routing branch from bypassing semantic features and making predictions on its own. RoTN captures the temporal evolution of anomalies from onset to persistence and termination. Experiments on three benchmarks show that RoMod achieves state-of-the-art performance while running substantially faster than dense backbones of comparable size.
Figures & tables
Figure 1: Observations on a frozen MoE backbone on XD-Violence test set. (a) Comparison of readout APs: native answer ( 71.37% ), hidden probe ( 82.87% ), and a single expert’s activation count ( 77.40% ). (b) Layer-wise expert activation contrast between normal and abnormal videos (red: high, green: low frequency).
Figure 2: Sparse expert specialization and input-dependent routing. (a) Per-expert discriminability from activation counts alone, with no hidden features or trainable parameters. (b) Routing preference over L20–L39 relative to the global average, grouped by event type; cells are comparable horizontally across sub-panels.
Figure 3: Overall framework of RoMod. A single forward pass yields the semantic hidden state from the optimal layer l⋆ and holistic routing statistics across all L layers; RoMF turns the routing vector into per-channel scales that recalibrate the centred hidden state, and RoTN emits clip scores incrementally under causal attention.
UBnormal
UCF-Crime
XD-Violence
Method
Venue
Data
AUC (%)
AUC (%)
AUC (%)
AP (%)
UR-DMU [ Zhou et al., 2023 ]
AAAI’23
100%
–
86.97
94.02
81.66
VadCLIP [ Wu et al., 2024 ]
AAAI’24
100%
–
88.02
–
84.51
VERA [ Ye et al., 2025 ]
CVPR’25
100%
–
86.55
88.26
70.54
π -VAD [ Majhi et al., 2025 ]
CVPR’25
100%
–
90.33
–
85.37
RefineVAD [ Lee et al., 2026 ]
AAAI’26
100%
–
88.92
–
88.66
Table 1: Comparison on three benchmarks. Data denotes the fraction of training videos used.
Table 5
Backbone
Type
Act. (B)
FLOPs (T)
RoMF (ms)
RoTN (ms)
Thpt. (FPS)
UCF-Crime AUC (%)
XD-Violence AP (%)
Qwen3.6-35B-A3B
Sparse
2.438
5.39
0.322
1.093
51.83
88.92
90.50
Qwen3VL-30B-A3B
Sparse
2.850
6.12
0.318
1.085
36.50
87.23
88.33
Gemma-4-26B-A4B
Sparse
3.684
10.37
0.325
0.839
39.56
86.79
87.24
Qwen3.8-27B
Dense
24.353
54.13
—
—
23.19
83.45
85.12
Table 4: Cost and streaming performance in online mode. Activated parameters (Act.) and FLOPs are measured per step; for the dense backbone Act. counts all non-embedding parameters.
Figure 4: Qualitative behaviour and layer-wise gains. (a) Frame-level scores of RoMod (blue) and Hidden-only (orange) on three videos; pink bands are ground-truth windows. (b, c) Layer-wise AP on XD-Violence and AUC on UCF-Crime.
Appendix figures & tables20 assets
Supplementary material from the paper’s appendix.
Appendix
Symbol
Where
Meaning
Architecture and indices
L , l
everywhere
number of MoE decoder layers; layer index. Lmoe is used when comparing backbones with different depths.
E , e
everywhere
number of routed experts per layer; expert index.
Ke
everywhere
number of experts activated per token (Top- Ke gating), Ke=8 .
D , c
everywhere
hidden width; hidden-channel index.
T , t
everywhere
number of clips in a video; clip index.
Appendix
Table 5: Symbol conventions. Letters that carry more than one role are disambiguated by subscript or by decoration; the table records which role applies where.
Seed Index
UBnormal
UCF-Crime
XD-Violence
AUC (%)
AUC (%)
AUC (%)
AP (%)
Run 1 (Seed 42, default)
79.17
91.32
97.34
92.49
Run 2 (Seed 1024)
78.36
90.64
96.85
91.82
Run 3 (Seed 2024)
78.85
91.03
97.12
92.15
Run 4 (Seed 3407)
77.42
89.65
96.08
90.73
Run 5 (Seed 7777)
78.11
90.28
96.53
91.36
Appendix
Table 6: Multi-seed evaluation across 5 independent runs on three benchmarks using 5% training data. Seed 42 is the default split reported in the main paper; all other runs fall within a ∼1.8 -point range of it.
Figure 5: Statistical hypothesis testing of category-dependent routing specialization on XD-Violence. (a) Layer-averaged pairwise Jensen–Shannon divergence matrix across six anomaly types. Every pair yields p^=0.0005 under 2,000 non-parametric permutation tests (marked ∗∗∗ ), which is the resolution bound of the test. (b) Directional Kullback–Leibler divergence matrix KL(Row∥Col) , quantifying the asymmetric information loss when approximating one category’s routing preference by another.
Figure 6: Channel–expert alignment on the default backbone (Qwen3.6-35B-A3B, XD-Violence). (a) Channel-level regression scatter at Layer 40 ( D=2048 ). Green rings mark the Top- 50 high-discriminability channels, which form a wedge envelope concentrated at large expert driving contributions Φc ( r=0.589 , asymptotic p<10−10 ). (b) Layer-wise alignment across all 40 MoE layers: Pearson r and Spearman ρ (top), Top- 50 channel overlap count and hypergeometric enrichment significance −log10penrich (bottom; the dash-dotted line marks αBonf=1.25×10−3 ). Both metrics sit at chance level throughout L1–L18 and rise sharply from L19 onward.
Figure 7: Cross-architecture verification on Gemma-4-26B-A4B (XD-Violence). (a) Channel-level scatter at Layer 30 ( D=2816 ). Despite fewer experts and a wider hidden state, the concentration of discriminative channels at large Φc replicates ( r=0.259 , asymptotic p<10−10 ). (b) Layer-wise alignment across 30 layers ( αBonf=1.67×10−3 ), showing the same transition: near-zero alignment in shallow layers, monotonic escalation through middle-to-late depths, and 16/30 layers reaching significance.
Backbone
Lmoe
D
E
rˉ (all)
rˉ (upper)
rmax (layer)
Sig. layers ( p<αBonf )
Peak Top-50 overlap (vs. chance baseline)
Qwen3.6-35B-A3B
40
2048
256
0.230
0.410
0.589 (L40)
25 / 40
23 / 50 ( 19.0× , p≈10−26 )
Gemma-4-26B-A4B
30
2816
128
0.107
0.196
0.326 (L28)
16 / 30
12 / 50 ( 13.5× , p≈10−11 )
Appendix
Table 7: Cross-architecture summary of channel–expert alignment on XD-Violence. rˉ (all) is the mean Pearson correlation over all MoE layers; rˉ (upper) is the mean over the upper half of the network. Significant layers satisfy penrich<αBonf=0.05/Lmoe with r>0 . The chance overlap baseline is E[X]=Kch2/D with Kch=50 .
Intervention
Kko /layer
Masked
AP (%)
AUC (%)
Unperturbed (frozen LM-head)
0
0
71.37
85.62
Mask Top- Kko
1
20
69.85 ( −1.52 )
84.32 ( −1.30 )
5
100
67.42 ( −3.95 )
81.65 ( −3.97 )
10
200
65.58 ( −5.79 )
79.20 ( −6.42 )
20
400
64.12 ( −7.25 )
76.85 ( −8.77 )
Mask control (rank >20 )
1
20
71.18 ( −0.19 )
85.50 ( −0.12 )
Appendix
Table 8: Causal intervention via layer-wise expert knockout on XD-Violence (layers 20–39, 97,396 clips). Scores are read from the frozen language head with no fitted parameters, so the reference row equals the native-answer readout of Table 14 . Masking the Kko most anomaly-preferring experts degrades the readout disproportionately to the removed capacity, whereas masking an identical number of control experts (rank >20 ) leaves it largely intact. Interventions are applied at all 20 middle-to-late layers simultaneously.
Figure 8: Complete 40-layer Fisher separability profiles on the frozen Qwen3.6-35B-A3B backbone. J(l) (Eq. 12 ) is computed from the terminal-token hidden states hNseq(l) , with no training and no temporal context. Top: XD-Violence ( 73,814 normal / 23,582 abnormal clips), peak 0.6842 at L28, 6.3× gain, plateau L28–L34. Middle: UCF-Crime ( 42,746 / 3,664 ), peak 0.5459 at L33, 2.8× gain, plateau L28–L37. Bottom: UBnormal ( 1,711 / 2,201 ), peak 0.1378 at L28, 9.9× gain, plateau L28–L31. Blue shading marks the shallow-collapse region (L1–L17); the coloured band marks the high-plateau zone. Absolute scales differ by an order of magnitude across datasets, yet the shallow floor, the L18–L27 rise and the plateau onset at L28 are shared.
Dataset
Clips (normal / abn.)
Shallow collapse
Jˉ (L1–L17)
High plateau
Peak layer / Jpeak
Gain
Plateau width / Jˉ
XD-Violence
73,814 / 23,582
L1–L17
0.109
L28–L34
L28 / 0.6842
6.3×
7 / 0.661
UCF-Crime
42,746 / 3,664
L1–L17
0.195
L28–L37
L33 / 0.5459
2.8×
10 / 0.521
UBnormal
1,711 / 2,201
L1–L17
0.014
L28–L31
L28 / 0.1378
9.9×
4 / 0.128
Appendix
Table 9: Stage boundaries of the Fisher separability profiles in Figure 8 . The gain is Jpeak divided by the mean J over the shallow-collapse region. Plateau width counts layers whose J(l) stays within 10% of the peak.
Figure 9: Layer-wise separability criterion G(l) on XD-Violence under different supervision budgets. Curves show G(l) estimated from clip-level pseudo-labels induced on 1% , 5% and 10% of weakly labelled training videos, against the 100% ground-truth reference. These are pseudo-label budgets, not zero-shot readouts. Pseudo-label noise attenuates the absolute value but preserves the ordering: all curves share the shallow floor at G≈0.22 , the rise from L20, and the plateau from L28. Spearman rank correlation with the ground-truth curve over L20–L40 is 0.78 , 0.93 and 0.96 for the three budgets.
Selection signal
Test GT
Selected layer l⋆
UBnormal
UCF-Crime
XD-Viol.
UBn
UCF
XD
AUC
AUC
AP
Oracle (downstream argmax)
yes
29
34
31
79.31
91.44
92.61
Fisher J(l) (hidden space)
yes
28
33
28
79.17
91.32
92.49
G(l) , 1% pseudo-labels
no
28
31
27
78.52
90.95
92.10
G(l) , 5% pseudo-labels (default)
no
28
33
28
79.17
91.32
92.49
G(l) , 10% pseudo-labels
no
29
33
28
79.31
91.32
92.49
Appendix
Table 10: Adaptive versus oracle layer selection ( λreg=0.1 , offline mode). Oracle selection maximizes the downstream test metric per dataset and is an upper bound conditional on the default 5% Stage-II budget ; it requires test ground truth and is therefore not deployable. J(l) ranks layers using test-set temporal boundaries in hidden space only. G(l) rows are fully deployable, using pseudo-labels induced from the stated fraction of weakly labelled training videos. In every row the Stage-II supervision budget is held at 5% , so the rows isolate the effect of layer selection and are not comparable to Table 3 , which varies the budget end-to-end. Dataset sub-columns follow the same left-to-right order on both sides of the table.
λreg
l⋆
Region
Feature alignment
Thpt. (FPS)
XD-AP (%)
Δ AP vs. λreg=0
0.00
33
Plateau
Fully developed
51.83
90.70
–
0.05
31
Plateau
Fully developed
51.83
90.63
−0.07
0.10
28
Plateau (onset)
Fully developed
51.83
90.50
−0.20
0.20
26
Emergence
Partially established
51.83
89.14
−1.56
0.40
21
Emergence
Nascent / incomplete
51.83
85.42
−5.28
0.80
20
lmin bound
Negligible alignment
51.83
84.61
−6.09
Appendix
Table 11: Sensitivity of the depth regularization weight λreg on XD-Violence (online streaming mode, 5% supervision budget). Throughput is unchanged across settings because all MoE layers contribute to the global routing fingerprint Rclip regardless of which layer supplies the semantic state. λreg=0.1 is the default operating point.
Dataset
l⋆
AUC (%)
AP (%)
Pos. base rate (%)
UBnormal
28
74.31±0.79
79.28±0.63
56.13
UCF-Crime
33
87.16±0.48
45.82±0.76
0 7.91
XD-Violence
28
94.72±0.31
88.14±0.42
24.24
Appendix
Table 12: Routing-only linear discriminant baseline. A linear head is fitted on Rclip(t) alone, without hidden states, temporal context or the backbone language head. Evaluation layers match the adaptive selection of Section 4.1 ; values are mean ± std over 5 random seeds. AUC is frame-level ROC-AUC.
Operator
Form
UBnormal
UCF-Crime
XD-Violence
AUC (%)
AUC (%)
AP (%)
Single-branch references
Hidden only
Hˉ
72.40
82.90
84.87
Routing only
WR
74.31
87.16
88.14
Fusion operators
Direct concatenation
[Hˉ;WR]
74.26
87.32
88.63
Appendix
Table 13: Comparison of fusion operators. All rows use one fusion block without RoTN, at the adaptively selected layer ( l⋆=28/33/28 for UBnormal / UCF-Crime / XD-Violence). The two reference rows are the single-branch limits: hidden-only is the ×/× row of Table 3 , routing-only is Table 12 . Metrics follow the main text: AUC for UBnormal and UCF-Crime, AP for XD-Violence. Best per column in bold.
Backbone model
Activated / total params
XD-Violence
UCF-Crime
UBnormal
AP (%)
AUC (%)
AUC (%)
AUC (%)
Random guess
—
24.24
50.00
50.00
50.00
Qwen3.6-35B-A3B
∼ 2.4B / 35B
71.37
85.62
74.95
58.27
Qwen3VL-30B-A3B
∼ 2.8B / 30B
76.65
89.02
75.83
59.41
Gemma-4-26B-A4B
∼ 3.7B / 26B
71.11
86.33
73.80
57.14
Appendix
Table 14: Zero-shot performance of frozen MoE backbones across three benchmarks without parameter tuning, read from native language output logits on visual clips.
λs
AUC (%)
AP (%)
TV ( ↓ )
Mean aˉ
0.05
97.05
90.89
0.0156
0.349
0.10
97.34
92.49
0.0212
0.319
0.50
96.94
91.01
0.0198
0.372
1.00
97.21
92.01
0.0213
0.317
2.00
96.77
90.73
0.0128
0.361
5.00
96.38
89.66
0.0102
0.401
Appendix
Table 15: Hyperparameter sensitivity sweep over λs and λp on XD-Violence ( l⋆=28 , averaged over 3 random seeds). The controlled variable is fixed at its operating value ( λp=0.05 in panel (a), λs=0.1 in panel (b)). Best AP per panel in bold. The peak matches the full model of Table 1 ( 97.34% AUC / 92.49% AP).
Component
Hyperparameter
Setting / value
Data sampling
Weak supervision ratio
5% of weakly labelled training videos
Frames per clip ( Tf )
24 frames
Sampled frames per clip ( K )
4 uniform frames
Input image dimensions
336×336 RGB (Lanczos interpolation)
RoMF module
Semantic hidden dimension ( D )
2048 (inherited from MLLM backbone)
Global routing dimension ( L⋅E )
10,240 ( 40 layers ×256 experts)
Appendix
Table 16: Hyperparameter specifications for Stage-II RoMod training and temporal modelling. All values correspond to the default configuration used for every reported result.
Multimodal Large Language Models (MLLMs) provide a natural way to make video anomaly detection more explainable. However, their final decisions do not always fully use the discriminative information contained in their hidden states, an issue we refer to as representation--behavior misalignment. We decompose this gap into a capacity component that measures discriminative information never aggregated into the readout position, and a directional component that measures the angular mismatch between the optimal and the native normal--abnormal axis at that position. Across multiple video anomaly detection benchmarks and MLLM backbones the directional component dominates, and residual-stream tracing shows that native-axis separability rises sharply in several mid-to-late attention layers. Because both components are governed by attention rather than MLP updates, we propose Representation--Behavior Alignment (RBA), a parameter-efficient method that adapts those layers using video-level labels alone while updating about 0.012% of the backbone parameters. Experiments on three benchmarks show that RBA improves native-readout performance and better aligns the model's decision direction with discriminative representations, and it produces anomaly decisions and explanations through a single generative process.
Existing Video Anomaly Detection (VAD) methods typically rely on task-specific training, leading to strong domain dependency and high training costs. Moreover, most existing methods output only scalar anomaly scores, providing limited insight into why specific events are considered abnormal. Recent advances in Vision-Language Models (VLMs) have enabled both anomaly detection and human-interpretable reasoning. However, many VLM-based approaches still require additional training steps (e.g., instruction tuning or verbalized learning) or external Large Language Models (LLMs), incurring further training costs and inference overhead. To address these challenges, we propose CoReVAD, a contextual reasoning framework for training-free video anomaly detection that operates with a single frozen VLM. CoReVAD directly generates anomaly scores and temporal descriptions from the VLM. To mitigate noise in generative outputs, we introduce a Local Response Cleaning (LRC) module based on local vision-text alignment. Furthermore, global temporal context and progression are incorporated through softmax-based refinement, Gaussian smoothing, and position weighting. Experiments on UCF-Crime and XD-Violence demonstrate that CoReVAD achieves competitive performance among training-free methods while providing reliable and interpretable explanations. Our official code is available at: https://github.com/Muk-00/CoReVAD
Hyeongmuk Lim, Youngbum Hur
Department of Industrial Engineering, Inha University, Incheon, Republic of Korea
Video anomaly detection (VAD) with multimodal large language models has shown strong potential, yet most existing methods still depend on large-scale annotations or expert-designed priors, limiting their ability to acquire anomaly knowledge with as little human intervention as possible. To address this, we propose Linguistic Relative Policy Optimization (LRPO), which distills group-relative semantic advantages from multiple reasoning trajectories into a linguistically expressed anomaly experience prior, and adapts the model by injecting this prior into the context to steer its output distribution without any parameter updates. LRPO builds two complementary experience representations: general experience captures transferable anomaly preferences across scenarios, while scenario experience models context-dependent anomaly rules for targeted refinement. To further improve the learned experience, we introduce an anomaly alignment reward that guides trajectory optimization to match human risk preferences and reinforce temporally grounded reasoning. Extensive experiments on XD-Violence, UCF-Crime, and UBnormal demonstrate that LRPO significantly outperforms existing state-of-the-art methods under tuning-free settings.
Jiaxu Leng, Jiankang Zheng, Mengjingcheng Mo +4
School of Computer Science and Technology, Chongqing University of Posts and Telecommunications, Chongqing, China · Chongqing College of Artificial Intelligence, Chongqing, China