Multimodal mixture-of-experts (MoE) models combine sparse expert activation with visual-language capabilities, yet their inference remains costly because long visual-token sequences repeatedly incur attention, routing, dispatch, and expert-MLP computation. Existing methods typically compress either the token or expert dimension, leaving redundancy along the other. Our analysis reveals two complementary regularities: the depth required for visual propagation varies across inputs, while text-token routing exhibits concentrated and recurrent expert-importance patterns. Based on these observations, we propose DecoMoE, a two-dimensional structured compression framework that decouples visual propagation from expert computation. The Sample-Adaptive Visual Boundary (SAVB) predicts an input-dependent visual-exit layer at which the visual-token block is removed. The Routing-Calibrated Expert Prefix (RCEP) reorders experts offline using text-token routed mass and, from this predicted exit layer onward, retains at each MoE layer the shortest contiguous prefix covering a target routed-mass fraction. We evaluate DecoMoE on Qwen3-VL-MoE and InternVL3.5-30B-A3B across six benchmarks. On Qwen3-VL-MoE, DecoMoE retains 97.91% of dense-baseline performance while reducing computation from 27.06 to 16.73 TFLOPs and latency from 0.44 to 0.26 seconds, yielding a 1.69x speedup. Code will be available at https://github.com/ShawnTan86/DecoMoE.
Figures & tables
Figure 1: Computation characteristics and structured compression of multimodal MoE inference. (a) Per-layer latency composition across one prefill pass and one decoding pass, shown as two consecutive 48-layer segments. Except for a small number of boundary-layer outliers, Attention and the expert MLP jointly account for more than 80% of the measured inference latency. (b) Comparison of compression across the token and expert dimensions, corresponding to attention and expert-MLP costs. Token pruning and expert compression sparsify one dimension, whereas DecoMoE removes the visual-token block and retains a contiguous prefix of experts. Opaque cells denote routed activations; gray regions denote removed computation.
Figure 2: Overview of DecoMoE. Top: At each candidate-layer entrance, SAVB determines whether to remove the complete visual-token block before self-attention. Once triggered at layer l , RCEP is activated in the MoE block of the same layer and independently restricts the triggering layer and each subsequent MoE layer to a mass-sufficient contiguous prefix of offline-reordered experts. Bottom left: SAVB constructs its gate feature from the pooled preceding text states and the final text state, and is trained with labels derived from forced-exit correctness and first-token NLL while the backbone remains frozen. Bottom right: Offline, RCEP derives a layer-wise Mean-Mass expert order from cross-case text-token routing. Online, it aggregates the activated routed mass over tokens in each MoE invocation and retains the shortest prefix covering the target mass fraction p .
Figure 3: Behavioral diagnostics for decoupled compression. (a) Forced visual-exit perplexity trajectories for different cases. Each curve reports the perplexity of the reference answer when visual tokens are removed after different prefill layers, and each marker indicates the earliest prefill layer at which forced visual exit still produces the correct answer. (b) Dataset-level fractions of routed mass covered by the top 20% of experts for visual and text tokens across layers. (c) Cross-case proportions of repeated experts for visual and text tokens across layers under fixed expert-retention ratios of 1/8 , 1/4 , and 1/2 .
Method
MMBench
ScienceQA
AI2D
AOKVQA
MMMU
HallusionBench
Avg. Ret.
Avg. TFLOPs
Latency (s)
Qwen3-VL-30B-A3B
Baseline
88.98
90.90
87.21
87.60
58.26
74.97
100.00%
27.06
0.44 (1.00 × )
FastV
87.74
87.31
82.93
86.11
57.21
71.67
96.32%
16.65
0.36 (0.82 × )
SparseVLM
87.59
88.51
83.78
86.29
56.35
69.63
96.02%
16.21
0.37 (0.84 × )
DyVTE
60.97
87.97
82.74
87.07
52.53
74.21
91.63%
18.29
0.37 (0.84 × )
FastMMoE
82.18
85.33
74.61
80.17
54.35
67.09
90.35%
18.88
1.56 (3.56 × )
Table 1: Accuracy–efficiency comparison on Qwen3-VL-MoE and InternVL3.5-30B-A3B. Avg. Ret. is the mean performance retention across six benchmarks; Avg. TFLOPs is the theoretical computation per profiled case; Latency is mean profiler CUDA time. Parentheses report latency relative to the corresponding baseline.
Figure 4: Dataset-wise SAVB exit distributions on Qwen3-VL-MoE. Triggered cases are grouped by zero-based layer intervals; Cases labeled Not Triggered retain visual tokens through all layers.
Method
POPE
HallusionBench
AI2D
Avg. Ret.
Avg. Exit Layer
Avg. Prefix Size
Avg. TFLOPs
Latency (s)
Baseline
89.82
74.97
87.21
100.00%
–
128
27.04
0.37 (1.00 × )
RCEP only
88.52
69.51
84.84
96.18%
28.00
34.86
16.69
0.28 (0.76 × )
SAVB + Fixed Prefix
89.08
71.40
85.49
97.48%
28.07
50
17.17
0.29 (0.79 × )
SAVB only
89.06
70.95
85.85
97.41%
28.08
128
17.85
0.30 (0.81 × )
Full DecoMoE
89.27
71.92
85.62
97.83%
28.07
35.01
17.07
0.26 (0.69 × )
Table 2: Component ablation on Qwen3-VL-MoE over POPE, HallusionBench, and AI2D. RCEP only uses a fixed visual exit at layer L28 ; SAVB + Fixed Prefix retains 50 experts after exit; SAVB only keeps all 128 experts.
Order Mode
Keep Ratio
10%
20%
30%
40%
50%
Native
87.3%
90.9%
97.0%
96.7%
97.4%
Random
85.4%
89.8%
94.9%
95.5%
95.9%
RCEP
96.9%
97.4%
97.7%
97.9%
97.5%
Table 3: RCEP analysis on Qwen3-VL-MoE over POPE, HallusionBench, and AI2D. (a) Performance retention under different expert orders and fixed keep ratios. (b) Sensitivity to the routed-mass target p , where Avg. Prefix denotes the mean retained expert-prefix length.
Figure 5: Analysis of DecoMoE’s visual boundary and realized acceleration. (a) Validation of the visual-propagation boundary predicted by SAVB, where visual exit is forced at offsets ranging from 8 layers before to 8 layers after the predicted boundary. (b) Latency–performance trade-off among representative efficient inference methods on Qwen3-VL-MoE over AI2D and HallusionBench; higher retention and lower latency lie toward the upper-left region.
Method
Baseline
FastV
SparseVLM
DyVTE
FastMMoE
DecoMoE
TTFT (ms)
370.80
320.77
339.13
358.93
1515.57
256.53
TPOT (ms/token)
135.80
117.48
124.20
131.45
555.05
93.95
Table 4: TTFT and TPOT comparison on Qwen3-VL-MoE. Lower is better.
Appendix figures & tables8 assets
Supplementary material from the paper’s appendix.
Appendix
Order Mode
Keep Ratio
10%
15%
20%
25%
30%
40%
50%
Native
87.26%
86.60%
90.85%
95.39%
96.97%
96.67%
97.41%
Random
85.43%
94.77%
89.79%
91.57%
94.93%
95.50%
95.94%
RCEP
96.87%
97.39%
97.41%
97.55%
97.71%
97.91%
97.47%
Appendix
Table 5: Complete mass–performance comparison of Native, Random, and RCEP expert orders on Qwen3-VL-MoE over POPE, HallusionBench, and AI2D. Entries report Avg. Ret. under fixed expert keep ratios.
Metric
Mass Target p
0.1
0.2
0.3
0.4
0.5
0.6
0.7
0.8
0.9
1.0
Avg. Ret.
96.7%
96.8%
96.7%
96.9%
97.2%
97.6%
97.8%
97.2%
97.6%
97.3%
Avg. Prefix
8.0
8.0
8.7
11.7
17.0
24.7
35.0
49.0
67.0
128.0
Appendix
Table 6: Complete sensitivity sweep for the RCEP routed-mass target p on Qwen3-VL-MoE over POPE, HallusionBench, and AI2D. Avg. Prefix is the mean retained expert-prefix length.
Figure 6: Token-to-expert routing map. Visual-token routing is diffuse and spans a broad expert set, whereas text-token routing forms a substantially more concentrated and recurrent structure.
Figure 7: Text-token router distributions at layer 40 for three cases, before and after visual-token eviction. The two conditions retain nearly identical structures.
Figure 8: Complementary routed-mass diagnostics. Top: number of experts required to cover 50% routed mass; most text cases require roughly 50–80 experts, whereas visual tokens lie near 120–128 experts. Bottom: deep-layer text-routing stability after visual-token eviction; cosine similarity remains above 0.997 and MAE below 0.0003.
Figure 9: The snowy residential scene used in Case 1. Its high-resolution input produces approximately 4,000 visual tokens.
Figure 10: The compact OCR input used in Case 2, containing approximately 80 visual tokens.
Figure 11: The fire-hydrant image shared by the three prompts in Case 3.
Institute of Automation, Chinese Academy of Sciences · School of Artificial Intelligence, University of Chinese Academy of Sciences · 3Zhongguancun Academy