Organizations: EIT-NLP Lab, Eastern Institute of Technology, Ningbo · Shanghai Jiao Tong University · The Hong Kong Polytechnic University · Munich Center for Machine Learning, LMU Munich
Multimodal large language models (MLLMs) incur substantial inference costs when processing long visual-textual sequences. While existing operation compression methods exploit modality-level redundancy, they largely treat computation within attention heads and shared feed-forward network (FFN) channels as unified units, leaving finer-grained redundancy underexplored. We find that redundancy varies both across modality-interaction paths within the same attention head and across visual and textual executions of the same FFN channel. Based on these findings, we propose Modality-aware Width-wise Operation Pruning (MWOP), which independently prunes visual-to-visual (V2V), text-to-visual (T2V), and text-to-text (T2T) attention paths within each layer, and separately selects FFN channels for visual and textual inputs. A first-order Taylor criterion guides the pruning process, with FFN importance re-evaluated after attention pruning and LoRA-based recovery training. To translate the resulting fine-grained sparsity into practical acceleration, we further develop path-sparse Triton attention kernels and compact visual-side FFN execution. MWOP preserves the token sequence while reducing attention and FFN computation, making it complementary to token compression and enabling simultaneous reduction of sequence length and per-token computation. On LLaVA-OneVision-7B, MWOP alone achieves a 1.6× prefill speedup with 99.7% average performance retention across 12 benchmarks. Combined with two representative token compression methods, it further increases their prefill speedups from 2.0× and 1.9× to 2.9× and 2.7×, respectively. Results on Qwen2.5-VL-7B further demonstrate its applicability across architectures. The code is available at https://github.com/EIT-NLP/MWOP.
Figures & tables
Figure 1: Comparison of operation compression methods. (a) ShortV skips visual token updates in selected layers. (b) RedundancyLens sparsifies visual-side attention and FFN computation. (c) MWOP performs finer-grained width compression by independently pruning V2V, T2V, and T2T attention paths and modality-conditioned FFN channel executions while preserving the token sequence. (d) MWOP achieves a 1.6 × prefill speedup with negligible performance loss. Combining it with token compression increases the speedup to 2.9 × .
Figure 2: Overview of MWOP. (1) MWOP performs width pruning while preserving the token sequence and remaining compatible with token pruning. (2) V2V , T2V , and T2T paths are masked independently within each attention head. (3) Shared FFN channels are selectively executed for visual tokens, textual tokens, or both.
Figure 3: Path-wise attention redundancy. Rows show General (top) and Grounding (bottom). (a) Overlap of the bottom-50% Taylor-ranked head sets across paths. (b) Single-path masking with the other paths kept dense (left) and whole-head versus head-wise path-level masking using Taylor or random selection at matched numbers of masked path blocks (right). All analyses are training-free, with performance retention measured relative to the original dense LLaVA-OneVision-7B model.
Setting
Mask Ratio (%)
Aggregate (%)
Category (%)
V2V
T2V
T2T
ATR
GRR
OGR
Gen.
Reas.
OCR
Grd.
Dense
0
0
0
100.0
100.0
100.0
100.0
100.0
100.0
100.0
V2V
30
0
0
100.1
100.1
100.1
100.3
99.9
100.1
100.1
40
0
0
100.0
100.3
99.7
100.6
100.1
99.4
99.9
40R
0
0
96.0
98.7
93.4
99.8
97.6
95.3
91.5
50
0
0
99.7
100.3
99.2
100.5
100.0
98.9
99.5
Table 1: Attention-path pruning with post-training recovery. Setting indicates the path whose masking ratio is varied at each stage. Scores are normalized to the dense reference. ATR reports overall retention. GRR covers General and Reasoning, while OGR covers OCR and Grounding. Superscript R denotes uniformly random block selection for the marked path, while unmarked paths use Taylor selection. Taylor and random counterparts share the same masking ratios and recovery protocol. Bold and underlined values indicate the best and second-best results within each stage, respectively. Gray shading marks the selected configurations. Per-benchmark results are provided in Appendix C.1 .
Figure 4: Modality-conditioned FFN redundancy. (a) Taylor (solid) vs. random (dashed) selection for All, Vision, and Text masking, with retention averaged equally across four task families. (b) Shared (Global) vs. modality-specific (Vision+Text) channel masking at matched per-modality ratios (left) and sensitivity to visual-only vs. text-only masking (right). All analyses are training-free, with performance measured relative to the attention-pruned model MWOPAttn with dense FFNs.
Setting
Mask Ratio (%)
Aggregate (%)
Category (%)
Visual
Text
ATR
GRR
OGR
Gen.
Reas.
OCR
Grd.
MWOP Attn
0
0
100.0
99.9
100.1
100.3
99.5
99.9
100.3
Visual Neuron
40
0
100.0
99.9
100.0
100.1
99.7
99.7
100.3
40R
0
96.4
97.3
95.6
99.0
95.6
94.3
96.9
50
0
99.7
99.9
99.6
100.4
99.3
98.9
100.2
50R
0
94.9
96.3
93.6
98.5
94.2
92.7
94.4
Table 2: Modality-conditioned FFN pruning with post-training recovery. Visual and Text denote FFN masking ratios. Scores are normalized to the dense reference. Superscript R denotes uniformly random channel selection for the marked modality, while unmarked masks use Taylor selection. Taylor and random counterparts share the same masking ratios and recovery protocol. Per-benchmark results are provided in Appendix C.1 .
Method
Sparsity
Ret. (%)
General
Reasoning
OCR
Grounding
ATR (%)
GRR (%)
OGR (%)
Dep.
Wid.
Tok.
GQA
VQA v2
VQA ok
SQA
AI2D
MMStar
VQA T
VQA D
OCRB
Ref
Ref+
RefG
LLaVA-OneVision-7B
-
-
-
100.0
63.1
82.2
59.9
96.0
81.9
61.4
75.3
87.4
63.5
77.9
70.5
76.9
100.0
100.0
100.0
ShortV
✓
65.6
62.5
81.0
58.9
94.7
79.8
60.1
63.6
70.7
49.5
46.3
42.1
42.8
84.0
98.3
69.7
RedundancyLens
✓
58.4
62.5
81.3
58.9
92.9
80.2
58.1
71.1
81.9
59.9
62.6
54.8
60.8
92.1
97.6
86.6
HalfV
✓
✓
60.4
62.6
81.4
59.7
95.6
81.4
59.9
73.1
85.1
55.0
52.7
46.4
47.8
89.2
99.1
79.4
DOP V
✓
✓
✓
58.3
62.6
80.7
58.7
95.3
81.4
61.2
67.0
73.2
54.1
54.8
46.9
52.2
88.0
98.9
77.1
Table 3: Comparison with operation compression methods on LLaVA-OneVision-7B at approximately 60% retained LLM-side FLOPs. Dep., Wid., and Tok. denote depth, width, and token compression; some compared methods also compress tokens. ATR, GRR, and OGR measure performance retention relative to the dense model. Unshaded rows apply compression after post-training; red and blue rows use post-training recovery with compression applied, with blue denoting ours. † marks methods originally proposed without post-compression recovery but evaluated with it here. MWOP achieves the best performance with the fewest FLOPs among compared methods, retaining 99.7% overall (ATR) and 99.6% fine-grained performance (OGR).
Method
Sparsity
Ret. (%)
General
Reasoning
OCR
Grounding
ATR (%)
GRR (%)
OGR (%)
Dep.
Wid.
Tok.
GQA
VQA v2
VQA ok
SQA
AI2D
MMStar
VQA T
VQA D
OCRB
Ref
Ref+
RefG
LLaVA-OneVision-7B
-
-
-
100.0
63.1
82.2
59.9
96.0
81.9
61.4
75.3
87.4
63.5
77.9
70.5
76.9
100.0
100.0
100.0
FastV
✓
32.8
59.9
77.6
58.6
89.7
77.0
54.6
61.1
60.4
44.6
47.5
39.5
46.5
80.1
93.9
66.3
PyramidDrop ‡
✓
33.4
60.5
80.1
59.8
90.2
76.2
52.1
70.2
75.2
40.5
16.0
13.0
13.7
72.1
94.2
50.0
VisionZip
✓
32.7
61.8
80.3
59.8
91.5
76.8
55.8
69.7
55.2
50.1
8.5
7.4
9.4
70.3
95.9
44.7
DART
✓
32.8
62.2
80.9
59.3
92.0
78.0
57.1
71.7
75.6
54.9
22.3
19.4
20.8
77.6
96.7
58.6
Table 4: Comparison of token and operation compression methods on LLaVA-OneVision-7B at approximately 30% retained LLM-side FLOPs. MWOP P and MWOP Z combine MWOP with PyramidDrop and ZOO-Prune. † denotes originally training-free methods evaluated with post-training recovery, while ‡ denotes originally training-based methods evaluated without post-training recovery. By composing operation and token compression, MWOP Z and MWOP P outperform all baselines overall while using fewer FLOPs.
Method
Theoretical (TFLOPs)
Practical (Prefill Latency, ms)
Attn
FFN
Total
Attn
FFN
Total
LLaVA-OneVision-7B
Dense
7.42
36.89
44.31
45.8
188.3
237.0
w/o Vis. Token
0.03
0.22
0.25
2.5 ( 18.1× )
9.7 ( 19.4× )
13.2 ( 18.0× )
w/o V2V&T2V
5.32
36.89
42.21
32.7 ( 1.4× )
188.3 ( 1.0× )
221.5 ( 1.1× )
w/o FFN V
7.42
0.22
7.63
45.8 ( 1.0× )
10.8 ( 17.4× )
56.9 ( 4.2× )
Table 5: Efficiency and cross-model generalization. Left : Theoretical FLOPs and practical prefill latency with breakdowns on two models. The component-removal configurations serve as references for potential speedup. Combining MWOP with token compression further reduces FLOPs and prefill latency on both architectures. Right : Cross-model generalization on Qwen2.5-VL-7B at approximately 40% retained LLM-side FLOPs, with standalone MWOP at 67.4% included as a reference. Composition improves overall and fine-grained retention over token compression alone. Per-benchmark results are reported in Appendix Table 13 .
Figure 5: Task-specialized attention-path compression. Bars show path compression ratios, and markers show post-recovery performance retention.
Settings
Multi-Image
Video
Avg (%)
Q-Bench2
Mantis-Eval
BLINK
Avg (%)
MVBench
Video-MME
NExT-QA
w/o Sub.
w/ Sub.
LLaVA-OneVision-7B
100.0
72.6
54.8
48.8
100.0
57.9
58.7
62.0
78.7
ZOO-Prune †
100.8
75.0
54.4
48.7
98.9
56.3
58.0
61.4
79.2
PyramidDrop
90.6
70.9
46.5
43.6
96.9
54.9
56.7
60.0
78.2
ShortV †
85.3
60.4
46.1
43.2
92.1
52.5
52.4
57.1
75.9
Table 6: Multi-image and video evaluation on LLaVA-OneVision-7B. All methods are evaluated directly after post-training recovery on single-image data. Sub. denotes subtitles.
Appendix figures & tables22 assets
Supplementary material from the paper’s appendix.
Appendix
Figure 6: Comparison of attention-path importance criteria on LLaVA-OneVision-7B. Rows correspond to V2V , T2V , and T2T masking. Columns respectively show GQA, MMStar, TextVQA, and RefCOCOg, representing General, Reasoning, OCR, and Grounding. Curves use Attention Mass, Attention Contribution, or first-order Taylor importance for training-free path selection.
Figure 7: Layer-head distribution of attention-path Taylor importance on LLaVA-OneVision-7B. Panels show V2V , T2V , and T2T . Rows and columns index decoder layers and query heads, using one-based labels. Darker cells indicate larger importance scores, and red boxes highlight the depth ranges discussed in the text.
Figure 8: Overlap of low-importance attention paths across four task categories. Each circle contains the layer-head identities in the bottom 50% of Taylor scores for one path: 392 of 784 identities on LLaVA-OneVision-7B. Region labels give the numbers of identities in the corresponding exclusive or shared sets. The central region contains heads whose three paths all lie in their respective low-importance sets.
Figure 9: Whole-head versus path-wise attention masking on LLaVA-OneVision-7B. Both use Taylor importance. The x-axis gives the fraction of masked heads or the independently masked fraction of each path type, matching the total number of removed path blocks up to rounding. Panels show training-free performance retention on the four task categories relative to the dense reference. The dashed line denotes 100% retention.
Figure 10: Training-free sensitivity to single-path attention masking on LLaVA-OneVision-7B. In each curve, Taylor-guided masking is applied only to the indicated path, with the other two paths kept dense. The horizontal axis gives the masking ratio of that path. The vertical axis gives performance retention relative to the dense reference. Panels cover the four task categories, and the dashed line denotes 100% retention.
Figure 11: Cross-task similarity of attention-path Taylor importance on LLaVA-OneVision-7B. Each matrix compares the layer-head importance profiles of General, Reasoning, OCR, and Grounding for one path. Cell values report pairwise Spearman rank correlations, with larger values indicating more similar rankings.
Setting
Aggregate (%)
Category (%)
ATR
GRR
OGR
Gen.
Reas.
OCR
Grd.
Dense
100.0
100.0
100.0
100.0
100.0
100.0
100.0
All Tasks
100.0
99.9
100.1
100.3
99.5
99.9
100.3
General
99.5
99.4
99.6
100.2
98.6
98.9
100.2
Reasoning
96.6
98.6
94.5
97.9
99.3
90.3
98.8
OCR
99.9
99.1
100.7
99.8
98.5
100.6
100.8
Appendix
Table 7: Effect of the calibration task on attention-path pruning. Performance retention (%) is measured relative to the unpruned, post-trained LLaVA-OneVision-7B.
Figure 12: Comparison of FFN-channel importance criteria on LLaVA-OneVision-7B. Columns show GQA, MMStar, TextVQA, and RefCOCOg for General, Reasoning, OCR, and Grounding. The top and bottom rows respectively mask visual-side and text-side channels. Random curves average three seeds.
Figure 13: Layer-wise agreement of FFN Taylor rankings across model states. Spearman rank correlation compares channel rankings from the same initial checkpoint before and after applying the selected attention mask, with scores estimated from all tokens, visual tokens, or text tokens.
Ratio
Setting
General
Reasoning
OCR
Grounding
ATR (%)
GRR (%)
OGR (%)
GQA
VQA v2
VQA ok
SQA
AI2D
MMStar
VQA T
VQA D
OCRB
Ref
Ref+
RefG
0%
Dense
63.1
82.2
59.9
96.0
81.9
61.4
75.3
87.4
63.5
77.9
70.5
76.9
100.0
100.0
100.0
10%
Old
63.2
81.8
59.5
95.4
81.0
60.1
73.6
85.1
66.6
77.8
69.9
76.4
99.4
99.2
99.7
New
64.4
83.6
56.9
96.9
82.3
63.6
74.8
85.3
65.7
77.1
69.8
76.2
100.1 (↑0.7)
100.6 (↑1.5)
99.6 (↓0.1)
20%
Old
63.0
81.6
58.2
94.1
79.8
59.1
72.4
84.0
64.9
76.7
69.0
75.5
98.1
98.0
98.2
New
64.0
82.8
55.7
97.0
82.3
63.8
73.6
85.8
65.7
77.2
69.9
75.2
99.7 (↑1.6)
100.1 (↑2.1)
99.2 (↑1.1)
Appendix
Table 8: Effect of recomputing FFN rankings under attention masking. Dense is the unmasked post-trained reference used to normalize ATR, GRR, and OGR. Bold marks the better Old/New result at each masking ratio. Arrows indicate changes from Old to New in percentage points.
Figure 14: Shared versus modality-specific FFN masks across four task categories. All curves use Taylor-guided selection on LLaVA-OneVision-7B under the fixed (40%,60%,10%) attention-path mask and report training-free performance retention relative to the attention-only reference.
Figure 15: Sensitivity to modality-specific FFN masking. Vision FFN and Text FFN mask visual and text token channel executions respectively, while keeping the other modality dense. Masks are selected by Taylor importance on LLaVA-OneVision-7B with the fixed (40%,60%,10%) attention-path mask.
Setting
Path Mask (%)
General
Reasoning
OCR
Grounding
ATR (%)
GRR (%)
OGR (%)
V2V
T2V
T2T
GQA
VQA v2
VQA ok
SQA
AI2D
MMStar
VQA T
VQA D
OCRB
Ref
Ref+
RefG
Dense Model
0
0
0
63.1
82.2
59.9
96.0
81.9
61.4
75.3
87.4
63.5
77.9
70.5
76.9
100.0
100.0
100.0
V2V
30
0
0
63.2
82.2
60.4
96.4
81.6
61.2
75.2
87.1
64.0
78.5
70.2
77.0
100.1
100.1
100.1
40
0
0
63.2
82.2
60.9
96.2
81.6
61.7
74.5
86.7
63.7
78.0
70.4
76.8
100.0
100.3
99.7
40R
0
0
62.8
81.6
60.3
94.7
79.9
59.3
73.1
80.2
61.7
71.9
63.5
71.0
96.0
98.7
93.4
50
0
0
63.2
82.2
60.8
96.2
81.5
61.6
74.5
85.6
63.4
77.6
70.0
76.7
99.7
100.3
99.2
Appendix
Table 9: Per-benchmark attention-path masking results on LLaVA-OneVision-7B after recovery. Setting identifies the path varied at each stage, and the three mask columns specify the complete configuration.
Setting
Channel Mask (%)
General
Reasoning
OCR
Grounding
ATR (%)
GRR (%)
OGR (%)
Visual
Text
GQA
VQA v2
VQA ok
SQA
AI2D
MMStar
VQA T
VQA D
OCRB
Ref
Ref+
RefG
MWOP Attn
0
0
63.3
82.0
60.4
96.5
81.5
60.5
74.0
85.3
65.9
78.1
71.0
77.0
100.0
99.9
100.1
Visual Neuron
40
0
63.4
82.0
60.0
96.2
81.3
61.3
74.1
85.2
65.6
77.8
70.8
77.3
100.0
99.9
100.0
40R
0
63.1
81.0
58.9
92.3
79.4
57.5
70.9
80.4
61.5
76.0
67.7
74.7
96.4
97.3
95.6
50
0
63.4
82.0
60.5
96.2
81.3
60.5
73.8
85.0
64.5
77.9
70.7
77.1
99.7
99.9
99.6
50R
0
62.9
80.5
58.6
91.2
78.7
56.2
70.3
77.4
61.2
74.4
66.0
72.5
94.9
96.3
93.6
Appendix
Table 10: Per-benchmark FFN masking results on LLaVA-OneVision-7B after recovery. Visual and Text specify modality-conditioned channel masking ratios under the fixed (40%,60%,10%) attention-path mask. MWOP Attn is the attention-only reference with dense FFNs.
Setting
Path Mask (%)
General
Reasoning
OCR
Grounding
ATR (%)
GRR (%)
OGR (%)
V2V
T2V
T2T
GQA
VQA v2
VQA ok
SQA
AI2D
MMStar
VQA T
VQA D
OCRB
Ref
Ref+
RefG
Dense Model
0
0
0
62.9
82.5
59.3
87.8
83.0
63.6
83.4
90.9
84.8
88.9
82.2
86.1
100.0
100.0
100.0
V2V
40
0
0
63.0
82.5
58.9
87.8
83.1
63.0
83.7
91.1
85.0
88.7
82.0
85.9
99.9
99.8
100.0
50
0
0
63.2
82.5
58.9
87.7
82.9
63.6
83.6
90.7
84.7
89.1
82.2
86.1
100.0
99.9
100.0
60
0
0
62.9
82.5
59.3
87.0
82.9
63.5
83.5
90.4
84.7
88.9
82.4
86.0
99.9
99.8
99.9
70
0
0
62.9
82.3
59.0
88.0
82.6
62.9
83.0
89.5
84.9
88.4
81.5
85.7
99.5
99.7
99.4
Appendix
Table 11: Per-benchmark attention-path masking results on Qwen2.5-VL-7B after recovery. All path masks use Taylor importance. V2V is swept first, followed by T2V with V2V fixed at 60%, and then T2T with the first two ratios fixed at 60% and 70%, respectively.
Setting
Channel Mask (%)
General
Reasoning
OCR
Grounding
ATR (%)
GRR (%)
OGR (%)
Visual
Text
GQA
VQA v2
VQA ok
SQA
AI2D
MMStar
VQA T
VQA D
OCRB
Ref
Ref+
RefG
MWOP Attn
0
0
63.4
82.2
58.7
87.5
83.8
63.0
83.0
89.3
84.9
88.9
82.0
86.1
99.8
99.9
99.6
Visual Neuron
30
0
63.6
82.3
58.7
87.9
83.5
62.8
82.6
89.1
84.3
88.6
81.9
86.0
99.6
99.9
99.3
40
0
63.4
82.3
58.5
87.5
83.0
63.3
81.9
88.7
85.0
88.6
81.8
85.9
99.5
99.7
99.2
50
0
63.4
82.1
57.7
88.3
83.3
62.9
81.0
88.5
83.4
89.0
82.2
86.1
99.2
99.6
98.8
60
0
63.3
82.0
57.3
88.0
82.8
62.5
80.0
88.4
83.9
88.8
82.1
85.9
98.9
99.2
98.6
Appendix
Table 12: Per-benchmark FFN masking results on Qwen2.5-VL-7B after recovery. Visual and Text give modality-conditioned channel masking ratios under the fixed (60%,70%,0%) attention-path mask.
Method
Sparsity
General
Reasoning
OCR
Grounding
ATR (%)
GRR (%)
OGR (%)
Dep.
Wid.
Tok.
GQA
VQA v2
VQA ok
SQA
AI2D
MMStar
VQA T
VQA D
OCRB
Ref
Ref+
RefG
Qwen2.5-VL-7B
-
-
-
62.9
82.5
59.3
87.8
83.0
63.6
83.4
90.9
84.8
88.9
82.2
86.1
100.0
100.0
100.0
PyramidDrop ‡
✓
53.9
76.4
35.3
82.4
72.1
44.4
78.7
65.4
23.5
31.7
26.3
27.2
65.1
81.4
48.9
ZOO-Prune
✓
59.7
79.6
34.3
84.9
76.8
55.8
57.2
64.0
57.0
81.3
73.7
76.4
83.5
87.7
79.4
ShortV
✓
54.0
71.1
31.6
81.5
68.1
46.7
26.1
57.7
47.1
55.1
47.9
48.7
66.7
78.9
54.5
RedundancyLens
✓
57.0
78.1
34.4
83.2
75.9
55.8
72.8
80.2
69.5
78.0
70.1
72.2
86.0
86.2
85.7
Appendix
Table 13: Comparison of compression methods on Qwen2.5-VL-7B at approximately 40% retained LLM-side FLOPs. Standalone MWOP at 67.4% retained FLOPs is included as a reference.
Figure 16: Input-length dependence of computational cost and acceleration. Blue and orange bars represent single-image and video inputs. Input lengths are averaged over 12 single-image benchmarks and three video benchmarks (MVBench, Video-MME, and NExT-QA). Within each model block, the upper row reports theoretical attention, FFN, and total decoder FLOPs in TFLOPs, while the lower row reports measured attention, FFN, and prefill speedups relative to Dense. Lower FLOPs and higher speedups indicate greater efficiency under the corresponding input-length setting.
Model
LLaVA-OneVision-7B
Qwen2.5-VL-7B
LLM Backbone
Qwen2-7B-Instruct
Qwen2.5-7B
Blocks
28
28
Query Heads
28
28
KV Heads
4
4
Head Dim
128
128
Hidden Dim
3584
3584
Appendix
Table 14: Language-model backbone configurations of LLaVA-OneVision-7B and Qwen2.5-VL-7B.
Figure 17: Layer-head attention-path masks for LLaVA-OneVision-7B. Rows and columns denote decoder layers and query heads, respectively, using one-based labels. Colored cells are retained, while white cells are masked. The masks remove 314/784 V2V paths (40.1%), 470/784 T2V paths (59.9%), and 78/784 T2T paths (9.9%). Each path is selected independently within a head.
Layer
Pruned FFN Channels
Layer
Pruned FFN Channels
Vision
Text
Vision
Text
0
16,473 (87.0%)
0 (0.0%)
14
980 (5.2%)
0 (0.0%)
1
14,690 (77.5%)
0 (0.0%)
15
2,533 (13.4%)
0 (0.0%)
2
14,722 (77.7%)
0 (0.0%)
16
7,231 (38.2%)
0 (0.0%)
3
12,337 (65.1%)
0 (0.0%)
17
9,414 (49.7%)
0 (0.0%)
4
10,083 (53.2%)
0 (0.0%)
18
12,317 (65.0%)
0 (0.0%)
Appendix
Table 15: Layer-wise FFN masking for the adopted LLaVA-OneVision-7B configuration. Layers are indexed from 0 to 27. Counts and percentages refer to masked modality-conditioned channel executions out of 18,944 channels per layer. The allocation masks 50% of visual-side channels globally across layers while retaining all text-side channels for the entire input sequence in every decoder layer.
Dataset
Count
Category
COCO [ 26 ]
103,909
General
VG [ 21 ]
55,504
Grounding
GQA [ 14 ]
42,777
Reasoning
OCRVQA [ 35 ]
27,346
General + OCR
ChartQA [ 33 ]
10,000
Reasoning + OCR
DVQA [ 16 ]
10,000
Reasoning + OCR
Appendix
Table 16: Composition of the post-training data mixture.
Recent multimodal large language models (MLLMs), such as Qwen2.5-VL and InternVL3, generate large numbers of vision tokens for high-resolution inputs, leading to substantial computational cost. Existing vision token pruning methods either depend on cross-modal attention and cannot prune before the prefill stage, or rely on diversity estimation with high computational overhead. We observe that attention scores from both vision and text tokens peak at modality separator tokens, suggesting that these separators bridge the two modalities. Based on this observation, we propose SepPrune, an efficient, training-free, plug-and-play pruning method that uses the separator token as a unified query to rank and select informative vision tokens. SepPrune reuses the LLM's built-in projection parameters and requires no architectural changes. Experiments on Qwen2.5-VL-7B show that SepPrune achieves state-of-the-art performance, retaining 96.3% of the original accuracy while removing 80.2% of vision tokens.
Yuchen Wang, Qihui Zhu, Yang Liu +2
MoE Key Lab of BIPC, NEL-BITA, University of Science and Technology of China, Hefei, China · ChangXin Memory Technologies, Hefei, China · APKL of BIIP, IAI, Hefei Comprehensive National Science Center, Hefei, China
In multimodal large language models (MLLMs), inference cost is largely dominated by the visual token prefix rather than the language backbone, making token reduction a key factor for improving efficiency. Existing approaches typically assign independent importance scores to visual tokens and retain a fixed number of top-ranked tokens, implicitly assuming token independence and a uniform compression ratio across inputs. In this work, we reformulate visual token pruning as a sequential decision-making process. Specifically, we introduce a pointer-style selection mechanism that iteratively chooses informative tokens, conditioning each decision on previously selected ones, and dynamically determines when to stop via a learned termination action. This enables joint optimization of both the selected subset and its size. To enable end-to-end training under standard language modeling objectives, we design a differentiable relaxation based on a variance-preserving noise interpolation scheme, allowing gradients to propagate through the discrete selection process. Extensive experiments on LLaVA-v1.5-7B and Qwen2.5-VL-7B demonstrate that our approach consistently outperforms fixed-ratio baselines across different compression levels. Under aggressive pruning that removes 88.9% of visual tokens, our method preserves 94.6% of the original accuracy while achieving a 1.88x speed-up in prefill latency.
Multimodal large language models (MLLMs) increasingly process long visual-token sequences, increasing the overall inference computation. Existing acceleration methods usually remove visual tokens or skip visual-token updates in entire layers, but these coarse strategies may discard fine-grained evidence or suppress useful operators together with redundant ones. In this paper, we study visual-token computation from an answer-observable perspective and find that late visual-token updates can remain large while having little effect on answer-token representations. Motivated by this answer-silent redundancy, we decompose each Transformer layer into attention and FFN operators and show that useful visual computation is often operator-dominant and layer-dependent. We propose an operator-level visual-token skipping framework that preserves the full visual-token sequence while selectively bypassing redundant attention, FFN, or both. Experiments across three MLLM architectures and 10 VQA benchmarks show that our method achieves strong efficiency-accuracy trade-offs, reducing \textbf{33.7%} TFLOPs on Qwen3-VL while retaining \textbf{99.5%} of the vanilla model performance.
Zhaoyang Luo, Runmin Dong, Miao Yang +4
Tsinghua Shenzhen International Graduate School, Shenzhen, China · Sun Yat-sen University, Zhuhai, China · Tsinghua University, Beijing, China +1