Multimodal mixture-of-experts (MoE) models combine sparse expert activation with visual-language capabilities, yet their inference remains costly because long visual-token sequences repeatedly incur attention, routing, dispatch, and expert-MLP computation. Existing methods typically compress either the token or expert dimension, leaving redundancy along the other. Our analysis reveals two complementary regularities: the depth required for visual propagation varies across inputs, while text-token routing exhibits concentrated and recurrent expert-importance patterns. Based on these observations, we propose DecoMoE, a two-dimensional structured compression framework that decouples visual propagation from expert computation. The Sample-Adaptive Visual Boundary (SAVB) predicts an input-dependent visual-exit layer at which the visual-token block is removed. The Routing-Calibrated Expert Prefix (RCEP) reorders experts offline using text-token routed mass and, from this predicted exit layer onward, retains at each MoE layer the shortest contiguous prefix covering a target routed-mass fraction. We evaluate DecoMoE on Qwen3-VL-MoE and InternVL3.5-30B-A3B across six benchmarks. On Qwen3-VL-MoE, DecoMoE retains 97.91% of dense-baseline performance while reducing computation from 27.06 to 16.73 TFLOPs and latency from 0.44 to 0.26 seconds, yielding a 1.69x speedup. Code will be available at https://github.com/ShawnTan86/DecoMoE.
Figures & tables
Figure 1: Computation characteristics and structured compression of multimodal MoE inference. (a) Per-layer latency composition across one prefill pass and one decoding pass, shown as two consecutive 48-layer segments. Except for a small number of boundary-layer outliers, Attention and the expert MLP jointly account for more than 80% of the measured inference latency. (b) Comparison of compression across the token and expert dimensions, corresponding to attention and expert-MLP costs. Token pruning and expert compression sparsify one dimension, whereas DecoMoE removes the visual-token block and retains a contiguous prefix of experts. Opaque cells denote routed activations; gray regions denote removed computation.
Figure 2: Overview of DecoMoE. Top: At each candidate-layer entrance, SAVB determines whether to remove the complete visual-token block before self-attention. Once triggered at layer l , RCEP is activated in the MoE block of the same layer and independently restricts the triggering layer and each subsequent MoE layer to a mass-sufficient contiguous prefix of offline-reordered experts. Bottom left: SAVB constructs its gate feature from the pooled preceding text states and the final text state, and is trained with labels derived from forced-exit correctness and first-token NLL while the backbone remains frozen. Bottom right: Offline, RCEP derives a layer-wise Mean-Mass expert order from cross-case text-token routing. Online, it aggregates the activated routed mass over tokens in each MoE invocation and retains the shortest prefix covering the target mass fraction p .
Figure 3: Behavioral diagnostics for decoupled compression. (a) Forced visual-exit perplexity trajectories for different cases. Each curve reports the perplexity of the reference answer when visual tokens are removed after different prefill layers, and each marker indicates the earliest prefill layer at which forced visual exit still produces the correct answer. (b) Dataset-level fractions of routed mass covered by the top 20% of experts for visual and text tokens across layers. (c) Cross-case proportions of repeated experts for visual and text tokens across layers under fixed expert-retention ratios of 1/8 , 1/4 , and 1/2 .
Method
MMBench
ScienceQA
AI2D
AOKVQA
MMMU
HallusionBench
Avg. Ret.
Avg. TFLOPs
Latency (s)
Qwen3-VL-30B-A3B
Baseline
88.98
90.90
87.21
87.60
58.26
74.97
100.00%
27.06
0.44 (1.00 × )
FastV
87.74
87.31
82.93
86.11
57.21
71.67
96.32%
16.65
0.36 (0.82 × )
SparseVLM
87.59
88.51
83.78
86.29
56.35
69.63
96.02%
16.21
0.37 (0.84 × )
DyVTE
60.97
87.97
82.74
87.07
52.53
74.21
91.63%
18.29
0.37 (0.84 × )
FastMMoE
82.18
85.33
74.61
80.17
54.35
67.09
90.35%
18.88
1.56 (3.56 × )
Table 1: Accuracy–efficiency comparison on Qwen3-VL-MoE and InternVL3.5-30B-A3B. Avg. Ret. is the mean performance retention across six benchmarks; Avg. TFLOPs is the theoretical computation per profiled case; Latency is mean profiler CUDA time. Parentheses report latency relative to the corresponding baseline.
Figure 4: Dataset-wise SAVB exit distributions on Qwen3-VL-MoE. Triggered cases are grouped by zero-based layer intervals; Cases labeled Not Triggered retain visual tokens through all layers.
Method
POPE
HallusionBench
AI2D
Avg. Ret.
Avg. Exit Layer
Avg. Prefix Size
Avg. TFLOPs
Latency (s)
Baseline
89.82
74.97
87.21
100.00%
–
128
27.04
0.37 (1.00 × )
RCEP only
88.52
69.51
84.84
96.18%
28.00
34.86
16.69
0.28 (0.76 × )
SAVB + Fixed Prefix
89.08
71.40
85.49
97.48%
28.07
50
17.17
0.29 (0.79 × )
SAVB only
89.06
70.95
85.85
97.41%
28.08
128
17.85
0.30 (0.81 × )
Full DecoMoE
89.27
71.92
85.62
97.83%
28.07
35.01
17.07
0.26 (0.69 × )
Table 2: Component ablation on Qwen3-VL-MoE over POPE, HallusionBench, and AI2D. RCEP only uses a fixed visual exit at layer L28 ; SAVB + Fixed Prefix retains 50 experts after exit; SAVB only keeps all 128 experts.
Order Mode
Keep Ratio
10%
20%
30%
40%
50%
Native
87.3%
90.9%
97.0%
96.7%
97.4%
Random
85.4%
89.8%
94.9%
95.5%
95.9%
RCEP
96.9%
97.4%
97.7%
97.9%
97.5%
Table 3: RCEP analysis on Qwen3-VL-MoE over POPE, HallusionBench, and AI2D. (a) Performance retention under different expert orders and fixed keep ratios. (b) Sensitivity to the routed-mass target p , where Avg. Prefix denotes the mean retained expert-prefix length.
Figure 5: Analysis of DecoMoE’s visual boundary and realized acceleration. (a) Validation of the visual-propagation boundary predicted by SAVB, where visual exit is forced at offsets ranging from 8 layers before to 8 layers after the predicted boundary. (b) Latency–performance trade-off among representative efficient inference methods on Qwen3-VL-MoE over AI2D and HallusionBench; higher retention and lower latency lie toward the upper-left region.
Method
Baseline
FastV
SparseVLM
DyVTE
FastMMoE
DecoMoE
TTFT (ms)
370.80
320.77
339.13
358.93
1515.57
256.53
TPOT (ms/token)
135.80
117.48
124.20
131.45
555.05
93.95
Table 4: TTFT and TPOT comparison on Qwen3-VL-MoE. Lower is better.
Appendix figures & tables8 assets
Supplementary material from the paper’s appendix.
Appendix
Order Mode
Keep Ratio
10%
15%
20%
25%
30%
40%
50%
Native
87.26%
86.60%
90.85%
95.39%
96.97%
96.67%
97.41%
Random
85.43%
94.77%
89.79%
91.57%
94.93%
95.50%
95.94%
RCEP
96.87%
97.39%
97.41%
97.55%
97.71%
97.91%
97.47%
Appendix
Table 5: Complete mass–performance comparison of Native, Random, and RCEP expert orders on Qwen3-VL-MoE over POPE, HallusionBench, and AI2D. Entries report Avg. Ret. under fixed expert keep ratios.
Metric
Mass Target p
0.1
0.2
0.3
0.4
0.5
0.6
0.7
0.8
0.9
1.0
Avg. Ret.
96.7%
96.8%
96.7%
96.9%
97.2%
97.6%
97.8%
97.2%
97.6%
97.3%
Avg. Prefix
8.0
8.0
8.7
11.7
17.0
24.7
35.0
49.0
67.0
128.0
Appendix
Table 6: Complete sensitivity sweep for the RCEP routed-mass target p on Qwen3-VL-MoE over POPE, HallusionBench, and AI2D. Avg. Prefix is the mean retained expert-prefix length.
Figure 6: Token-to-expert routing map. Visual-token routing is diffuse and spans a broad expert set, whereas text-token routing forms a substantially more concentrated and recurrent structure.
Figure 7: Text-token router distributions at layer 40 for three cases, before and after visual-token eviction. The two conditions retain nearly identical structures.
Figure 8: Complementary routed-mass diagnostics. Top: number of experts required to cover 50% routed mass; most text cases require roughly 50–80 experts, whereas visual tokens lie near 120–128 experts. Bottom: deep-layer text-routing stability after visual-token eviction; cosine similarity remains above 0.997 and MAE below 0.0003.
Figure 9: The snowy residential scene used in Case 1. Its high-resolution input produces approximately 4,000 visual tokens.
Figure 10: The compact OCR input used in Case 2, containing approximately 80 visual tokens.
Figure 11: The fire-hydrant image shared by the three prompts in Case 3.
Large-scale vision-language mixture-of-experts (VL-MoE) models provide strong multimodal capability, but efficient deployment on memory-constrained platforms remains difficult. Existing MoE offloading systems are largely designed for text-centric workloads and become much less effective for visual-heavy inputs, where large numbers of visual tokens induce broader and less predictable expert accesses. We present VisMMoE, a VL-MoE offloading system built on a single systems insight: pruning redundant visual tokens can improve offloading not only by reducing computation, but also by reshaping expert demand. We refer to this effect as \textit{visual-expert affinity}: token pruning makes expert accesses more concentrated within layers and more stable across layers, producing a smaller and more predictable expert working set. Guided by this insight, VisMMoE combines affinity-aware token compression, lookahead expert prediction, and cache/pipeline orchestration to improve expert locality and prefetch effectiveness under tight memory budgets. We implement VisMMoE on multiple frameworks and evaluate it on representative VL-MoE models and benchmarks. VisMMoE improves end-to-end inference performance by up to 2.68x and 1.61x, respectively, over strong baselines for today's VL-MoE deployments while maintaining competitive accuracy.
Cheng Xu, Xiaofeng Hou, Jiacheng Liu +1
Shanghai Jiao Tong University Minhang Qu, Shanghai, China
Mixture-of-Experts Multimodal Large Language Models (MoE-MLLMs) offer remarkable performance but incur prohibitive GPU memory costs, making compression essential. Among PTQ methods, expert-level mixed-precision quantization has proven effective for MoE-LLMs, yet suffers notable degradation on MoE-MLLMs due to two overlooked biases in expert importance estimation. (1) At the cross-modal level, the numerical dominance of vision tokens causes expert selection frequency to be dominated by vision tokens, masking experts that are critical to the text modality; (2) at the intra-vision level, the large proportion of redundant vision tokens further skew frequency statistics, obscuring experts critical for informative visual content. To bridge gaps, we propose MODE, a modality-decomposed expert-level mixed-precision quantization framework for MoE-MLLMs that decomposes expert selection frequency by modality, filters redundant vision tokens to obtain denoised visual frequency, and further evaluates quantization sensitivity per modality as a complementary signal to frequency-based estimation. These signals are integrated into an Integer Linear Programming formulation to assign per-expert bit-widths under a given budget. Extensive experiments show that MODE is particularly well-suited for MoE-MLLMs, limiting average performance loss to within 2.9% at W3A16, with larger gains at the extreme 2-bit setting.
Yuanteng Chen, Peisong Wang, Zhilei Liu +9
Institute of Automation, Chinese Academy of Sciences · School of Artificial Intelligence, University of Chinese Academy of Sciences · 3Zhongguancun Academy
Mixture-of-Experts (MoE) has become a prevalent backbone for large vision-language models (VLMs), yet how modality-specific signals should guide expert routing remains under-explored. Existing routing strategies are either hand-crafted or modality-agnostic, relying on idealized priors that ignore the layer-dependent modality fusion patterns in MoE-VLMs and provide little guidance for expert specialization. We propose Soft Modality-guided Expert Specialization (SMoES), which consists of dynamic soft modality scores that capture layer-dependent fusion patterns, an expert binning mechanism aligned with expert-parallel deployment, and an inter-bin mutual information regularization that encourages coherent modality specialization. Our method leverages attention-based or Gaussian-statistics modality scores to optimize mutual information regularization. Experiments across four MoE-based VLMs and 16 benchmarks demonstrate improvement on both effectiveness and efficiency: 0.9% and 4.2% average gain on multimodal and language-only tasks, 56.1% reduction in EP communication overhead, and 12.3% throughput improvement under realistic deployment. These results validate that aligning routing with modality-aware expert specialization unlocks MoE-VLM capacity and efficiency.