Long visual token sequences often account for a substantial fraction of the computational overhead in multimodal large language models~(MLLMs). Existing approaches reduce this cost by pruning redundant visual tokens, but permanently discard visual evidence that may become useful in subsequent layers. We instead ask whether all visual tokens can be preserved while reducing the cost of repeatedly evolving the representations through the Transformer. To answer this question, we perform low-rank interventions on visual-to-text information flow. We find that, after visual-to-text attention is blocked, restoring only a few directions recovers most of the lost accuracy, suggesting the relevant visual influence is concentrated in a low-dimensional subspace. We further observe strong predictability in layer-specific visual states: lightweight MLPs approximate them with high cosine similarity and low reconstruction error. Motivated by these findings, we propose δ-Vision, which replaces repeated Transformer evolution of visual tokens with lightweight low-rank adapters that construct layer-wise visual memories while preserving all visual tokens for text retrieval. Across image and video benchmarks, δ-Vision achieves higher accuracy than visual token pruning baselines at comparable or lower computation, while delivering competitive inference efficiency without discarding visual tokens.
Figures & tables
Figure 1 : Analysis of the dimensional structure of visual representations. (a) Downstream accuracy under different retained ranks of visual representation. (b) Layer-wise reconstruction quality measured by cosine similarity and reconstruction MSE.
Figure 2 : Overview of δ-Vision . Top: δ-Vision replaces the original visual hidden-state evolution with lightweight low-rank adapters that constructs layer-wise visual representations. Bottom left: we consider two visual-memory variants: Embedding Adapter , which predicts visual states from the initial visual embeddings, and Recurrent Adapter , which propagates a compact visual state across layers. Bottom right: δ-Vision is trained with Supervised-KD, where the frozen vanilla MLLM serves as the teacher and provides token-level distributional supervision under teacher forcing.
Method
MMStar
RWQA
GQA
MMB
MMB-CN
MME
POPE
SQA
VQA-v2
Avg.
Upper Bound (100% Retention)
Qwen3-VL-4B
64.9
71.2
61.6
87.5
87.7
84.7
89.3
93.3
80.9
80.1
Training-Free
Retain 20% Visual Tokens
FastV (ECCV24)
49.5
55.0
47.3
79.2
77.8
76.6
77.6
83.2
65.8
68.0
VisionZip (CVPR25)
51.1
61.7
52.9
82.0
80.8
78.6
82.6
85.2
70.8
71.7
Table 1 : Results of visual token compression methods under 20% and 5% visual-token retention.
Method
FLOPs
Peak Mem.
Total
Prefill
(%)
(GB)
Speedup
Speedup
Upper Bound
Qwen3-VL-4B
100.00
12.94
1.00
1.00
Retain 5% Visual Tokens
FastV
21.49
11.26
1.22
1.41
DART
21.23
11.12
1.13
1.24
Table 2 : Efficiency comparison on Video-MME under 5% visual-token retention. Total Speedup measures the overall inference speedup, including both prefill and decoding.
Model
Method
MMStar
RWQA
GQA
MMB
MMB-CN
MME
POPE
SQA
VQA-v2
Avg.
LLaVA-1.5-7B
Vanilla
37.0
54.9
61.2
72.5
67.9
71.7
83.4
65.6
75.9
65.6
DART
29.3
46.5
51.6
63.8
55.7
68.9
68.1
69.0
61.5
57.2
DivPrune
31.3
49.0
53.4
66.7
57.3
66.3
79.3
64.5
70.5
59.8
δ-Vision
35.4
51.4
53.8
69.9
63.7
68.7
79.2
65.4
69.5
61.9
Qwen3-VL-30B-A3B
Vanilla
70.3
72.3
65.2
89.4
90.8
89.4
87.9
96.3
82.6
82.7
DART
46.2
52.3
50.1
77.2
77.9
79.2
74.3
85.9
66.7
67.8
Table 3 : Results of different visual token compression methods across various MLLMs. Vanilla denotes the uncompressed upper bound. DART and DivPrune retain 5% of the visual tokens.
Method
MuirBench
Video-MME
MVBench
Avg.
Upper Bound
Qwen3-VL-4B
53.5
52.0
61.6
55.7
Training-Free (Retain 5% Visual Tokens)
FastV
42.1
45.3
41.2
42.9
DART
45.1
51.3
53.3
49.9
VisionZip
43.6
48.2
46.5
46.1
Table 4 : Results on multi-image and video benchmarks. Training-free baselines are evaluated at 5% visual-token retention.
Dataset
Q
QK⊤
Attention Output
Full
r95
ER
Full
r95
ER
Full
r95
ER
Qwen3-VL-4B
MMStar
213.9
42.7
108.4
119.2
5.0
19.3
212.2
47.0
101.2
RWQA
1296.0
112.7
530.0
128.0
10.7
35.3
1296.0
103.5
378.5
SQA
170.5
35.4
89.1
101.3
4.1
16.0
170.5
38.0
83.4
LLaVA-1.5-7B
Table 5 : Average rank statistics of visual-token computations across Transformer layers. r95 denotes the minimum rank retaining 95% of the squared singular-value energy, and ER denotes effective rank.
Figure 3 : Accuracy recovery under low-rank visual attention effect. We vary the retained rank r from 0 to 1024.
Appendix figures & tables8 assets
Supplementary material from the paper’s appendix.
Appendix
Method
MMStar
RWQA
GQA
MMB
MMB-CN
MME
POPE
SQA
VQA-v2
Avg.
Upper Bound (100% Retention)
Qwen3-VL-4B
64.9
71.2
61.6
87.5
87.7
84.7
89.3
93.3
80.9
80.1
Training-Free
Retain 15% Visual Tokens
FastV
47.5
50.5
44.1
76.7
74.8
75.1
73.1
80.3
63.7
65.1
VisionZip
49.7
58.4
53.4
80.2
78.4
76.4
79.5
82.3
70.3
69.8
Appendix
Table 7 : Results of visual token compression methods on Qwen3-VL-4B. Training-free baselines are evaluated under 15% and 10% visual-token retention.
Method
MuirBench
Video-MME
MVBench
Avg.
Upper Bound
Qwen3-VL-4B
53.5
52.0
61.6
55.7
Training-Free (Retain 20% Visual Tokens)
FastV
48.6
49.5
54.6
50.9
DART
49.0
52.2
56.7
52.6
VisionZip
47.9
50.9
55.0
51.3
Appendix
Table 8 : Additional results on multi-image and video benchmarks. Training-free baselines are evaluated at 20% visual-token retention.
Method
FLOPs
Peak Mem.
Total
Prefill
Decode
(%)
(GB)
Speedup
Speedup
Speedup
Upper Bound
Qwen3-VL-4B
100.00
12.94
1.00
1.00
1.00
Training-Free (Retain 20% Visual Tokens)
FastV
33.27
11.68
1.18
1.34
1.07
DART
33.00
11.64
1.13
1.23
1.06
Appendix
Table 9 : Efficiency comparison on Video-MME. Training-free baselines are evaluated under 20% visual-token retention. Total Speedup measures the overall inference speedup, including both prefill and decoding.
Method
MMStar
RWQA
GQA
MMB
MMB-CN
MME
POPE
SQA
VQA-v2
Avg.
Time
Emb. Adapter (Init).
41.2
40.8
43.4
77.4
74.6
66.4
65.5
78.8
58.4
60.7
-
Emb. Adapter + SFT
50.5
59.4
28.6
83.6
80.0
77.5
85.5
79.9
75.1
68.9
23m51s
Emb. Adapter + OPD
53.6
63.9
55.8
83.4
83.1
79.5
86.7
81.9
77.8
74.0
4h38m08s
Emb. Adapter + Supervised-KD
54.4
62.9
56.7
84.5
82.8
80.1
87.6
82.3
78.0
74.4
34m53s
Appendix
Table 10 : Ablation study on different training objectives for the embedding adapter. All variants only optimize the adapter parameters. Emb. Adapter (Init). denotes the initialization setting, where the same visual representation is shared across all Transformer layers without learned layer-specific adaptation, serving as the lower-bound baseline.
Model
Method
MMStar
RWQA
GQA
MMB
MMB-CN
MME
POPE
SQA
VQA-v2
Avg.
LLaVA-1.5-7B
Vanilla
37.0
54.9
61.2
72.5
67.9
71.7
83.4
65.6
75.9
65.6
DART
30.6
52.3
57.8
70.8
66.3
76.6
81.4
69.4
71.1
64.0
DivPrune
38.1
51.9
58.5
70.4
66.5
70.3
82.1
65.6
74.3
64.2
δ-Vision
35.4
51.4
53.8
69.9
63.7
68.7
79.2
65.4
69.5
61.9
LLaVA-1.5-13B
Vanilla
38.7
57.7
63.0
74.7
70.3
66.6
83.8
71.5
78.3
67.2
DART
33.6
55.4
57.6
76.0
67.9
75.6
84.9
73.6
72.1
66.3
Appendix
Table 11 : Results of different visual token compression methods across various vision-language models. Vanilla denotes the uncompressed upper bound. DART and DivPrune retain 20% of the visual tokens.
Model
Method
MMStar
RWQA
GQA
MMB
MMB-CN
MME
POPE
SQA
VQA-v2
Avg.
LLaVA-1.5-13B
Vanilla
38.7
57.7
63.0
74.7
70.3
66.6
83.8
71.5
78.3
67.2
DART
32.7
48.1
50.0
67.7
58.2
58.4
70.7
70.9
63.0
57.7
DivPrune
35.9
51.2
44.1
69.5
60.9
52.8
62.9
68.3
58.5
56.0
δ-Vision
36.5
50.1
56.3
72.5
65.9
64.2
76.8
70.7
73.1
62.9
LLaVA-1.6-Mistral-7B
Vanilla
40.4
47.3
54.3
71.9
65.0
68.4
85.4
68.7
74.4
64.0
DART
29.6
42.5
39.1
51.0
46.1
58.6
67.0
67.5
52.3
50.4
Appendix
Table 12 : Results of different visual token compression methods across various vision-language models. Vanilla denotes the uncompressed upper bound. DART and DivPrune retain 5% of the visual tokens.
Figure 4 : Performance of Embedding Adapter under different adapter ranks r . Dashed horizontal lines denote the corresponding vanilla model performance.
Figure 5 : Visual write and read interventions in the hybrid-attention backbone. Mask LA removes the visual write to the recurrent state in all Linear Attention (LA) layers by restoring the post-visual recurrent state to its pre-visual value. For Full Attention (FA) layers, we block visual read by preventing textual queries from attending to visual keys and values. We compare blocking FA layers 11 and 15 against blocking the other six FA layers (3, 7, 19, 23, 27, and 31).
Multimodal large language models (MLLMs) achieve strong performance across diverse vision-language tasks, but their efficiency is limited by the cost of processing numerous visual tokens. Visual token pruning can reduce this cost, but requires accurate token importance estimates. Recent studies have demonstrated that text-to-vision attention from middle language model layers can effectively guide visual token pruning, typically using attention from a predefined middle layer to select the visual tokens to retain. Two problems therefore remain. First, our analysis shows that the layer whose attention is most responsive to the question varies substantially across samples, making a fixed layer suboptimal. Second, obtaining attention from the appropriate middle layer requires processing numerous visual tokens through several language model layers, by which point considerable computation has already been spent. To address both problems, we propose Middle-layer Attention Prediction (MAP), which uses Question Contrastive Teacher Selection to identify a sample-specific teacher layer by contrasting attention under the original and reference questions, and distills attention from the selected layer into a lightweight predictor that estimates visual token importance from multi-modal input features. During inference, MAP combines the predicted importance scores with a diversity criterion to prune visual tokens before the first language model layer. Thus, MAP requires no attention maps for pruning and remains compatible with existing inference acceleration techniques. Across ten benchmarks on LLaVA-NeXT-7B, MAP retains 97.5% of the unpruned model performance with only 5.56% of the visual tokens, yielding a 3.09x end-to-end speedup.
Yuyao Sun, Tao Deng, Shuang Li +3
Beihang University · Shanghai EABOT Technology Co., Ltd.
Multimodal large language models (MLLMs) increasingly process long visual-token sequences, increasing the overall inference computation. Existing acceleration methods usually remove visual tokens or skip visual-token updates in entire layers, but these coarse strategies may discard fine-grained evidence or suppress useful operators together with redundant ones. In this paper, we study visual-token computation from an answer-observable perspective and find that late visual-token updates can remain large while having little effect on answer-token representations. Motivated by this answer-silent redundancy, we decompose each Transformer layer into attention and FFN operators and show that useful visual computation is often operator-dominant and layer-dependent. We propose an operator-level visual-token skipping framework that preserves the full visual-token sequence while selectively bypassing redundant attention, FFN, or both. Experiments across three MLLM architectures and 10 VQA benchmarks show that our method achieves strong efficiency-accuracy trade-offs, reducing \textbf{33.7%} TFLOPs on Qwen3-VL while retaining \textbf{99.5%} of the vanilla model performance.
Zhaoyang Luo, Runmin Dong, Miao Yang +4
Tsinghua Shenzhen International Graduate School, Shenzhen, China · Sun Yat-sen University, Zhuhai, China · Tsinghua University, Beijing, China +1
Modern multimodal large language models (MLLMs) typically keep the language model fixed and train a visual projector that maps the pixels into a sequence of tokens in its embedding space, so that images can be presented in essentially the same form as text. However, the language model has been optimized to operate on discrete, semantically meaningful tokens, while prevailing visual projectors transform an image into a long stream of continuous and highly correlated embeddings. This causes the visual tokens to behave differently from the word-like units that LLMs are originally trained to understand. We propose a novel Disentangled Visual Tokenization (DiVT) that clusters patch embeddings into coherent semantic units, so each token corresponds to a distinct visual concept instead of a rigid grid cell. DiVT further adapts its token budget to image complexity, providing an explicit accuracy-compute trade-off modifying neither the vision encoder nor the language model. Across diverse multimodal benchmarks, DiVT matches or surpasses baselines with significantly fewer visual tokens, demonstrating robustness under limited token budgets, significantly reducing memory cost and latency while making visual inputs more compatible with LLMs. Our code is available at https://github.com/snuviplab/DiVT.
Hyun Lee, Hyemin Jeong, Yejin Kim +4
Seoul National University · Ewha Womans University