MWOP: Modality-aware Width-wise Operation Pruning for Efficient MLLMs
Organizations: EIT-NLP Lab, Eastern Institute of Technology, Ningbo · Shanghai Jiao Tong University · The Hong Kong Polytechnic University · Munich Center for Machine Learning, LMU Munich
Abstract
Multimodal large language models (MLLMs) incur substantial inference costs when processing long visual-textual sequences. While existing operation compression methods exploit modality-level redundancy, they largely treat computation within attention heads and shared feed-forward network (FFN) channels as unified units, leaving finer-grained redundancy underexplored. We find that redundancy varies both across modality-interaction paths within the same attention head and across visual and textual executions of the same FFN channel. Based on these findings, we propose Modality-aware Width-wise Operation Pruning (MWOP), which independently prunes visual-to-visual (V2V), text-to-visual (T2V), and text-to-text (T2T) attention paths within each layer, and separately selects FFN channels for visual and textual inputs. A first-order Taylor criterion guides the pruning process, with FFN importance re-evaluated after attention pruning and LoRA-based recovery training. To translate the resulting fine-grained sparsity into practical acceleration, we further develop path-sparse Triton attention kernels and compact visual-side FFN execution. MWOP preserves the token sequence while reducing attention and FFN computation, making it complementary to token compression and enabling simultaneous reduction of sequence length and per-token computation. On LLaVA-OneVision-7B, MWOP alone achieves a prefill speedup with 99.7% average performance retention across 12 benchmarks. Combined with two representative token compression methods, it further increases their prefill speedups from and to and , respectively. Results on Qwen2.5-VL-7B further demonstrate its applicability across architectures. The code is available at https://github.com/EIT-NLP/MWOP.
Figures & tables
| Setting | Mask Ratio (%) | Aggregate (%) | Category (%) | |||||||
| V2V | T2V | T2T | ATR | GRR | OGR | Gen. | Reas. | OCR | Grd. | |
| Dense | 0 | 0 | 0 | 100.0 | 100.0 | 100.0 | 100.0 | 100.0 | 100.0 | 100.0 |
| V2V | 30 | 0 | 0 | 100.1 | 100.1 | 100.1 | 100.3 | 99.9 | 100.1 | 100.1 |
| 40 | 0 | 0 | 100.0 | 100.3 | 99.7 | 100.6 | 100.1 | 99.4 | 99.9 | |
| 0 | 0 | 96.0 | 98.7 | 93.4 | 99.8 | 97.6 | 95.3 | 91.5 | ||
| 50 | 0 | 0 | 99.7 | 100.3 | 99.2 | 100.5 | 100.0 | 98.9 | 99.5 | |
| Setting | Mask Ratio (%) | Aggregate (%) | Category (%) | ||||||
| Visual | Text | ATR | GRR | OGR | Gen. | Reas. | OCR | Grd. | |
| MWOP Attn | 0 | 0 | 100.0 | 99.9 | 100.1 | 100.3 | 99.5 | 99.9 | 100.3 |
| Visual Neuron | 40 | 0 | 100.0 | 99.9 | 100.0 | 100.1 | 99.7 | 99.7 | 100.3 |
| 0 | 96.4 | 97.3 | 95.6 | 99.0 | 95.6 | 94.3 | 96.9 | ||
| 50 | 0 | 99.7 | 99.9 | 99.6 | 100.4 | 99.3 | 98.9 | 100.2 | |
| 0 | 94.9 | 96.3 | 93.6 | 98.5 | 94.2 | 92.7 | 94.4 | ||
| Method | Sparsity | Ret. (%) | General | Reasoning | OCR | Grounding | ATR (%) | GRR (%) | OGR (%) | ||||||||||
| Dep. | Wid. | Tok. | GQA | VQA | VQA | SQA | AI2D | MMStar | VQA | VQA | OCRB | Ref | Ref+ | RefG | |||||
| LLaVA-OneVision-7B | - | - | - | 100.0 | 63.1 | 82.2 | 59.9 | 96.0 | 81.9 | 61.4 | 75.3 | 87.4 | 63.5 | 77.9 | 70.5 | 76.9 | 100.0 | 100.0 | 100.0 |
| ShortV | ✓ | 65.6 | 62.5 | 81.0 | 58.9 | 94.7 | 79.8 | 60.1 | 63.6 | 70.7 | 49.5 | 46.3 | 42.1 | 42.8 | 84.0 | 98.3 | 69.7 | ||
| RedundancyLens | ✓ | 58.4 | 62.5 | 81.3 | 58.9 | 92.9 | 80.2 | 58.1 | 71.1 | 81.9 | 59.9 | 62.6 | 54.8 | 60.8 | 92.1 | 97.6 | 86.6 | ||
| HalfV | ✓ | ✓ | 60.4 | 62.6 | 81.4 | 59.7 | 95.6 | 81.4 | 59.9 | 73.1 | 85.1 | 55.0 | 52.7 | 46.4 | 47.8 | 89.2 | 99.1 | 79.4 | |
| DOP V | ✓ | ✓ | ✓ | 58.3 | 62.6 | 80.7 | 58.7 | 95.3 | 81.4 | 61.2 | 67.0 | 73.2 | 54.1 | 54.8 | 46.9 | 52.2 | 88.0 | 98.9 | 77.1 |
| Method | Sparsity | Ret. (%) | General | Reasoning | OCR | Grounding | ATR (%) | GRR (%) | OGR (%) | ||||||||||
| Dep. | Wid. | Tok. | GQA | VQA | VQA | SQA | AI2D | MMStar | VQA | VQA | OCRB | Ref | Ref+ | RefG | |||||
| LLaVA-OneVision-7B | - | - | - | 100.0 | 63.1 | 82.2 | 59.9 | 96.0 | 81.9 | 61.4 | 75.3 | 87.4 | 63.5 | 77.9 | 70.5 | 76.9 | 100.0 | 100.0 | 100.0 |
| FastV | ✓ | 32.8 | 59.9 | 77.6 | 58.6 | 89.7 | 77.0 | 54.6 | 61.1 | 60.4 | 44.6 | 47.5 | 39.5 | 46.5 | 80.1 | 93.9 | 66.3 | ||
| PyramidDrop | ✓ | 33.4 | 60.5 | 80.1 | 59.8 | 90.2 | 76.2 | 52.1 | 70.2 | 75.2 | 40.5 | 16.0 | 13.0 | 13.7 | 72.1 | 94.2 | 50.0 | ||
| VisionZip | ✓ | 32.7 | 61.8 | 80.3 | 59.8 | 91.5 | 76.8 | 55.8 | 69.7 | 55.2 | 50.1 | 8.5 | 7.4 | 9.4 | 70.3 | 95.9 | 44.7 | ||
| DART | ✓ | 32.8 | 62.2 | 80.9 | 59.3 | 92.0 | 78.0 | 57.1 | 71.7 | 75.6 | 54.9 | 22.3 | 19.4 | 20.8 | 77.6 | 96.7 | 58.6 | ||
| Method | Theoretical (TFLOPs) | Practical (Prefill Latency, ms) | ||||
| Attn | FFN | Total | Attn | FFN | Total | |
| LLaVA-OneVision-7B | ||||||
| Dense | 7.42 | 36.89 | 44.31 | 45.8 | 188.3 | 237.0 |
| w/o Vis. Token | 0.03 | 0.22 | 0.25 | 2.5 ( ) | 9.7 ( ) | 13.2 ( ) |
| w/o V2V&T2V | 5.32 | 36.89 | 42.21 | 32.7 ( ) | 188.3 ( ) | 221.5 ( ) |
| w/o FFN V | 7.42 | 0.22 | 7.63 | 45.8 ( ) | 10.8 ( ) | 56.9 ( ) |
| Settings | Multi-Image | Video | |||||||
| Avg (%) | Q-Bench2 | Mantis-Eval | BLINK | Avg (%) | MVBench | Video-MME | NExT-QA | ||
| w/o Sub. | w/ Sub. | ||||||||
| LLaVA-OneVision-7B | 100.0 | 72.6 | 54.8 | 48.8 | 100.0 | 57.9 | 58.7 | 62.0 | 78.7 |
| ZOO-Prune | 100.8 | 75.0 | 54.4 | 48.7 | 98.9 | 56.3 | 58.0 | 61.4 | 79.2 |
| PyramidDrop | 90.6 | 70.9 | 46.5 | 43.6 | 96.9 | 54.9 | 56.7 | 60.0 | 78.2 |
| ShortV | 85.3 | 60.4 | 46.1 | 43.2 | 92.1 | 52.5 | 52.4 | 57.1 | 75.9 |
Appendix figures & tables22 assets
Supplementary material from the paper’s appendix.
Appendix
| Setting | Aggregate (%) | Category (%) | |||||
| ATR | GRR | OGR | Gen. | Reas. | OCR | Grd. | |
| Dense | 100.0 | 100.0 | 100.0 | 100.0 | 100.0 | 100.0 | 100.0 |
| All Tasks | 100.0 | 99.9 | 100.1 | 100.3 | 99.5 | 99.9 | 100.3 |
| General | 99.5 | 99.4 | 99.6 | 100.2 | 98.6 | 98.9 | 100.2 |
| Reasoning | 96.6 | 98.6 | 94.5 | 97.9 | 99.3 | 90.3 | 98.8 |
| OCR | 99.9 | 99.1 | 100.7 | 99.8 | 98.5 | 100.6 | 100.8 |
| Ratio | Setting | General | Reasoning | OCR | Grounding | ATR (%) | GRR (%) | OGR (%) | ||||||||
| GQA | VQA | VQA | SQA | AI2D | MMStar | VQA | VQA | OCRB | Ref | Ref+ | RefG | |||||
| 0% | Dense | 63.1 | 82.2 | 59.9 | 96.0 | 81.9 | 61.4 | 75.3 | 87.4 | 63.5 | 77.9 | 70.5 | 76.9 | 100.0 | 100.0 | 100.0 |
| 10% | Old | 63.2 | 81.8 | 59.5 | 95.4 | 81.0 | 60.1 | 73.6 | 85.1 | 66.6 | 77.8 | 69.9 | 76.4 | 99.4 | 99.2 | 99.7 |
| New | 64.4 | 83.6 | 56.9 | 96.9 | 82.3 | 63.6 | 74.8 | 85.3 | 65.7 | 77.1 | 69.8 | 76.2 | 100.1 | 100.6 | 99.6 | |
| 20% | Old | 63.0 | 81.6 | 58.2 | 94.1 | 79.8 | 59.1 | 72.4 | 84.0 | 64.9 | 76.7 | 69.0 | 75.5 | 98.1 | 98.0 | 98.2 |
| New | 64.0 | 82.8 | 55.7 | 97.0 | 82.3 | 63.8 | 73.6 | 85.8 | 65.7 | 77.2 | 69.9 | 75.2 | 99.7 | 100.1 | 99.2 | |
| Setting | Path Mask (%) | General | Reasoning | OCR | Grounding | ATR (%) | GRR (%) | OGR (%) | ||||||||||
| V2V | T2V | T2T | GQA | VQA | VQA | SQA | AI2D | MMStar | VQA | VQA | OCRB | Ref | Ref+ | RefG | ||||
| Dense Model | 0 | 0 | 0 | 63.1 | 82.2 | 59.9 | 96.0 | 81.9 | 61.4 | 75.3 | 87.4 | 63.5 | 77.9 | 70.5 | 76.9 | 100.0 | 100.0 | 100.0 |
| V2V | 30 | 0 | 0 | 63.2 | 82.2 | 60.4 | 96.4 | 81.6 | 61.2 | 75.2 | 87.1 | 64.0 | 78.5 | 70.2 | 77.0 | 100.1 | 100.1 | 100.1 |
| 40 | 0 | 0 | 63.2 | 82.2 | 60.9 | 96.2 | 81.6 | 61.7 | 74.5 | 86.7 | 63.7 | 78.0 | 70.4 | 76.8 | 100.0 | 100.3 | 99.7 | |
| 0 | 0 | 62.8 | 81.6 | 60.3 | 94.7 | 79.9 | 59.3 | 73.1 | 80.2 | 61.7 | 71.9 | 63.5 | 71.0 | 96.0 | 98.7 | 93.4 | ||
| 50 | 0 | 0 | 63.2 | 82.2 | 60.8 | 96.2 | 81.5 | 61.6 | 74.5 | 85.6 | 63.4 | 77.6 | 70.0 | 76.7 | 99.7 | 100.3 | 99.2 | |
| Setting | Channel Mask (%) | General | Reasoning | OCR | Grounding | ATR (%) | GRR (%) | OGR (%) | |||||||||
| Visual | Text | GQA | VQA | VQA | SQA | AI2D | MMStar | VQA | VQA | OCRB | Ref | Ref+ | RefG | ||||
| MWOP Attn | 0 | 0 | 63.3 | 82.0 | 60.4 | 96.5 | 81.5 | 60.5 | 74.0 | 85.3 | 65.9 | 78.1 | 71.0 | 77.0 | 100.0 | 99.9 | 100.1 |
| Visual Neuron | 40 | 0 | 63.4 | 82.0 | 60.0 | 96.2 | 81.3 | 61.3 | 74.1 | 85.2 | 65.6 | 77.8 | 70.8 | 77.3 | 100.0 | 99.9 | 100.0 |
| 0 | 63.1 | 81.0 | 58.9 | 92.3 | 79.4 | 57.5 | 70.9 | 80.4 | 61.5 | 76.0 | 67.7 | 74.7 | 96.4 | 97.3 | 95.6 | ||
| 50 | 0 | 63.4 | 82.0 | 60.5 | 96.2 | 81.3 | 60.5 | 73.8 | 85.0 | 64.5 | 77.9 | 70.7 | 77.1 | 99.7 | 99.9 | 99.6 | |
| 0 | 62.9 | 80.5 | 58.6 | 91.2 | 78.7 | 56.2 | 70.3 | 77.4 | 61.2 | 74.4 | 66.0 | 72.5 | 94.9 | 96.3 | 93.6 | ||
| Setting | Path Mask (%) | General | Reasoning | OCR | Grounding | ATR (%) | GRR (%) | OGR (%) | ||||||||||
| V2V | T2V | T2T | GQA | VQA | VQA | SQA | AI2D | MMStar | VQA | VQA | OCRB | Ref | Ref+ | RefG | ||||
| Dense Model | 0 | 0 | 0 | 62.9 | 82.5 | 59.3 | 87.8 | 83.0 | 63.6 | 83.4 | 90.9 | 84.8 | 88.9 | 82.2 | 86.1 | 100.0 | 100.0 | 100.0 |
| V2V | 40 | 0 | 0 | 63.0 | 82.5 | 58.9 | 87.8 | 83.1 | 63.0 | 83.7 | 91.1 | 85.0 | 88.7 | 82.0 | 85.9 | 99.9 | 99.8 | 100.0 |
| 50 | 0 | 0 | 63.2 | 82.5 | 58.9 | 87.7 | 82.9 | 63.6 | 83.6 | 90.7 | 84.7 | 89.1 | 82.2 | 86.1 | 100.0 | 99.9 | 100.0 | |
| 60 | 0 | 0 | 62.9 | 82.5 | 59.3 | 87.0 | 82.9 | 63.5 | 83.5 | 90.4 | 84.7 | 88.9 | 82.4 | 86.0 | 99.9 | 99.8 | 99.9 | |
| 70 | 0 | 0 | 62.9 | 82.3 | 59.0 | 88.0 | 82.6 | 62.9 | 83.0 | 89.5 | 84.9 | 88.4 | 81.5 | 85.7 | 99.5 | 99.7 | 99.4 | |
| Setting | Channel Mask (%) | General | Reasoning | OCR | Grounding | ATR (%) | GRR (%) | OGR (%) | |||||||||
| Visual | Text | GQA | VQA | VQA | SQA | AI2D | MMStar | VQA | VQA | OCRB | Ref | Ref+ | RefG | ||||
| MWOP Attn | 0 | 0 | 63.4 | 82.2 | 58.7 | 87.5 | 83.8 | 63.0 | 83.0 | 89.3 | 84.9 | 88.9 | 82.0 | 86.1 | 99.8 | 99.9 | 99.6 |
| Visual Neuron | 30 | 0 | 63.6 | 82.3 | 58.7 | 87.9 | 83.5 | 62.8 | 82.6 | 89.1 | 84.3 | 88.6 | 81.9 | 86.0 | 99.6 | 99.9 | 99.3 |
| 40 | 0 | 63.4 | 82.3 | 58.5 | 87.5 | 83.0 | 63.3 | 81.9 | 88.7 | 85.0 | 88.6 | 81.8 | 85.9 | 99.5 | 99.7 | 99.2 | |
| 50 | 0 | 63.4 | 82.1 | 57.7 | 88.3 | 83.3 | 62.9 | 81.0 | 88.5 | 83.4 | 89.0 | 82.2 | 86.1 | 99.2 | 99.6 | 98.8 | |
| 60 | 0 | 63.3 | 82.0 | 57.3 | 88.0 | 82.8 | 62.5 | 80.0 | 88.4 | 83.9 | 88.8 | 82.1 | 85.9 | 98.9 | 99.2 | 98.6 | |
| Method | Sparsity | General | Reasoning | OCR | Grounding | ATR (%) | GRR (%) | OGR (%) | ||||||||||
| Dep. | Wid. | Tok. | GQA | VQA | VQA | SQA | AI2D | MMStar | VQA | VQA | OCRB | Ref | Ref+ | RefG | ||||
| Qwen2.5-VL-7B | - | - | - | 62.9 | 82.5 | 59.3 | 87.8 | 83.0 | 63.6 | 83.4 | 90.9 | 84.8 | 88.9 | 82.2 | 86.1 | 100.0 | 100.0 | 100.0 |
| PyramidDrop | ✓ | 53.9 | 76.4 | 35.3 | 82.4 | 72.1 | 44.4 | 78.7 | 65.4 | 23.5 | 31.7 | 26.3 | 27.2 | 65.1 | 81.4 | 48.9 | ||
| ZOO-Prune | ✓ | 59.7 | 79.6 | 34.3 | 84.9 | 76.8 | 55.8 | 57.2 | 64.0 | 57.0 | 81.3 | 73.7 | 76.4 | 83.5 | 87.7 | 79.4 | ||
| ShortV | ✓ | 54.0 | 71.1 | 31.6 | 81.5 | 68.1 | 46.7 | 26.1 | 57.7 | 47.1 | 55.1 | 47.9 | 48.7 | 66.7 | 78.9 | 54.5 | ||
| RedundancyLens | ✓ | 57.0 | 78.1 | 34.4 | 83.2 | 75.9 | 55.8 | 72.8 | 80.2 | 69.5 | 78.0 | 70.1 | 72.2 | 86.0 | 86.2 | 85.7 | ||
| Model | LLaVA-OneVision-7B | Qwen2.5-VL-7B |
| LLM Backbone | Qwen2-7B-Instruct | Qwen2.5-7B |
| Blocks | 28 | 28 |
| Query Heads | 28 | 28 |
| KV Heads | 4 | 4 |
| Head Dim | 128 | 128 |
| Hidden Dim | 3584 | 3584 |
| Layer | Pruned FFN Channels | Layer | Pruned FFN Channels | ||
| Vision | Text | Vision | Text | ||
| 0 | 16,473 (87.0%) | 0 (0.0%) | 14 | 980 (5.2%) | 0 (0.0%) |
| 1 | 14,690 (77.5%) | 0 (0.0%) | 15 | 2,533 (13.4%) | 0 (0.0%) |
| 2 | 14,722 (77.7%) | 0 (0.0%) | 16 | 7,231 (38.2%) | 0 (0.0%) |
| 3 | 12,337 (65.1%) | 0 (0.0%) | 17 | 9,414 (49.7%) | 0 (0.0%) |
| 4 | 10,083 (53.2%) | 0 (0.0%) | 18 | 12,317 (65.0%) | 0 (0.0%) |
| Dataset | Count | Category |
| COCO [ 26 ] | 103,909 | General |
| VG [ 21 ] | 55,504 | Grounding |
| GQA [ 14 ] | 42,777 | Reasoning |
| OCRVQA [ 35 ] | 27,346 | General + OCR |
| ChartQA [ 33 ] | 10,000 | Reasoning + OCR |
| DVQA [ 16 ] | 10,000 | Reasoning + OCR |