Organizations: University of Science and Technology of China · Alibaba Token Hub, Alibaba Group · The Chinese University of Hong Kong · Zhejiang University
Multimodal large language models have become the dominant paradigm for visual understanding, but incur substantial costs by encoding inputs into dense, fixed-size patch tokens. However, visual information is unevenly distributed: some regions require fine-grained detail, while others admit compact representations. Downsampling sacrifices this detail, while existing token pruning and adaptive approaches remain limited in content-adaptive granularity, task generalization, and integration with modern MLLMs and serving infrastructure. Overcoming these limitations calls for foundation models that learn, end to end, where-and at what granularity-to allocate visual representations, a native capability we term elastic visual representation weaving. We introduce VisionWeave, establishing this capability in frontier-level MLLMs through large-scale training. It combines two components: a gated spatial pooler constructs coarse-grained representations alongside native fine-grained representations within a shared MRoPE coordinate, while a granularity router learns their content-adaptive allocation. Through self-distillation alone, we validate this capability on Qwen3.5-4B and scale to Qwen3.8-27B with over 30K A100 GPU-hours. Based on Qwen3.8-27B, VisionWeave adaptively adjusts token savings to visual content, saving 43.0% tokens on average while retaining 98.9% native performance across eight benchmarks, versus only 88% performance preserved for token pruning baselines with a fixed 50% savings target. Extensive evaluations confirm robust efficiency-quality trade-offs across diverse tasks, resolutions and video frames. When deployed on SGLang serving engine, our method achieves a 2.3x throughput gain while reducing mean TTFT by 54.4% and mean TPOT by 60.6%. Together, we believe these results position elastic visual weaving as a promising capability for next-generation multimodal models.
Figures & tables
Figure 1: Weaving elastic visual representations. More examples appear in Appendix C .
Figure 2: VisionWeave achieves content-adaptive token savings with near-native quality, whereas fix-budget baselines degrade sharply on challenging tasks. Visual case studies on ScreenSpotV2 and DocVQA illustrate this difference (Appendix B ). Our VisionWeave still retains higher quality even when baselines are granted its adaptive saving ratios (Appendix E.3 ).
Figure 3: Overall architecture of VisionWeave. The gated spatial pooler complements native fine-grained tokens Vfine with coarse-grained representations Vcoarse within the same MRoPE coordinate, while the granularity router adaptively weaves them into a mixed-granularity sequence for the LLM.
Stage 1
Stage 2
Stage 3
Trainable components
Gated Spatial Pooler
Granularity Router
Pooler, LLM
Teacher
Native model
Native model
Native model
Routing
Bypassed (all coarse)
Soft-mixing routing
Hard routing
Loss terms
Ldistill
Ldistill+0.02Lbal
Ldistill
Student visual tokens
25%
100%
Content-adaptive
Table 1: Three-stage self-distillation for elastic visual processing.
Source
Domain
Approx. training samples
LLaVA-OneVision ( An et al., 2025 )
Image understanding
246K
LLaVA-Video-178K ( Zhang et al., 2025b )
Video understanding
281K
UGround ( Gou et al., 2025 )
GUI grounding
250K
Total
777K
Table 2: Training-data composition for Qwen3.8-27B.
Setting
Stage 1
Stage 2
Stage 3
Peak learning rate
1×10−4
1×10−5
3×10−6
Minimum learning rate
5×10−6
0
0
Warmup fraction
5%
5%
2%
Global batch size
32
64
512
Training steps
5,257
2,500
1,000
Gradient clipping
1.0
1.0
0.5
Table 3: Training hyperparameters for Qwen3.8-27B.
Figure 4: Training dynamics of three-stage self-distillation.
Native Model
Fixed Savings (s=50%)
Content-adaptive Savings
Qwen3.5-4B
512 token/Image
Downsampling
FastV †
VisionZip
VisionWeave
Savings (%)
RealWorldQA
74.64
72.16 ( ↓ 3.32%)
70.98 ( ↓ 4.90%)
70.72 ( ↓ 5.25%)
73.46 ( ↓ 1.58%)
43.2
DocVQA
91.92
82.71 ( ↓ 10.02%)
78.69 ( ↓ 14.39%)
76.57 ( ↓ 16.70%)
91.03 ( ↓ 0.97%)
16.0
InfoVQA
67.44
52.87 ( ↓ 21.60%)
51.75 ( ↓ 23.27%)
50.89 ( ↓ 24.54%)
66.11 ( ↓ 1.97%)
10.6
ScreenSpotV2
89.86
80.42 ( ↓ 10.51%)
72.56 ( ↓ 19.25%)
70.44 ( ↓ 21.61%)
89.47 ( ↓ 0.43%)
24.1
HallusionBench
69.88
69.09 ( ↓ 1.13%)
69.26 ( ↓ 0.89%)
68.64 ( ↓ 1.77%)
69.26 ( ↓ 0.89%)
44.0
Table 4: Content-adaptive token savings with near-native quality.
Figure 5: VisionWeave achieves consistently favorable trade-offs across tasks and input resolutions.
Figure 10
Model
Throughput (req/min) ↑
Mean TTFT (s) ↓
P95 TTFT (s) ↓
Mean TPOT (ms) ↓
P95 TPOT (ms) ↓
Native Model
0.90
169.90
281.41
355.64
715.58
VisionWeave
2.07 ( 2.30× )
77.54 ( ↓ 54.4%)
120.59 ( ↓ 57.1%)
140.29 ( ↓ 60.6%)
289.19 ( ↓ 59.6%)
Table 5: VisionWeave accelerates Qwen3.8-27B serving on SGLang.
Figure 8: VisionWeave enables flexible inference-time control of token savings while closely matching native-model performance when savings are disabled.
Qwen3.5-4B
Qwen3.8-27B
Benchmark
VisionWeave
Shuffled
Δ
VisionWeave
Shuffled
Δ
RealWorldQA
73.46
73.46
0.00
77.91
76.60
−1.31
DocVQA
91.03
88.70
−2.32
91.33
88.35
−2.98
InfoVQA
66.11
64.83
−1.28
71.61
69.04
−2.58
ScreenSpotV2
89.47
80.03
−9.43
91.43
59.28
−32.15
HallusionBench
69.26
67.32
−1.95
73.60
71.66
−1.95
Table 6: Learned granularity allocation improves performance over shuffled routing.
Long visual token sequences often account for a substantial fraction of the computational overhead in multimodal large language models~(MLLMs). Existing approaches reduce this cost by pruning redundant visual tokens, but permanently discard visual evidence that may become useful in subsequent layers. We instead ask whether all visual tokens can be preserved while reducing the cost of repeatedly evolving the representations through the Transformer. To answer this question, we perform low-rank interventions on visual-to-text information flow. We find that, after visual-to-text attention is blocked, restoring only a few directions recovers most of the lost accuracy, suggesting the relevant visual influence is concentrated in a low-dimensional subspace. We further observe strong predictability in layer-specific visual states: lightweight MLPs approximate them with high cosine similarity and low reconstruction error. Motivated by these findings, we propose δ-Vision, which replaces repeated Transformer evolution of visual tokens with lightweight low-rank adapters that construct layer-wise visual memories while preserving all visual tokens for text retrieval. Across image and video benchmarks, δ-Vision achieves higher accuracy than visual token pruning baselines at comparable or lower computation, while delivering competitive inference efficiency without discarding visual tokens.
Jingdi lei, Junxian Li, Di Zhang +3
Nanyang Technological University · Shanghai Jiao Tong University · Fudan University +2
Recent multimodal large language models (MLLMs), such as Qwen2.5-VL and InternVL3, generate large numbers of vision tokens for high-resolution inputs, leading to substantial computational cost. Existing vision token pruning methods either depend on cross-modal attention and cannot prune before the prefill stage, or rely on diversity estimation with high computational overhead. We observe that attention scores from both vision and text tokens peak at modality separator tokens, suggesting that these separators bridge the two modalities. Based on this observation, we propose SepPrune, an efficient, training-free, plug-and-play pruning method that uses the separator token as a unified query to rank and select informative vision tokens. SepPrune reuses the LLM's built-in projection parameters and requires no architectural changes. Experiments on Qwen2.5-VL-7B show that SepPrune achieves state-of-the-art performance, retaining 96.3% of the original accuracy while removing 80.2% of vision tokens.
Yuchen Wang, Qihui Zhu, Yang Liu +2
MoE Key Lab of BIPC, NEL-BITA, University of Science and Technology of China, Hefei, China · ChangXin Memory Technologies, Hefei, China · APKL of BIIP, IAI, Hefei Comprehensive National Science Center, Hefei, China
Visual token pruning reduces the inference overhead of multimodal large language models (MLLMs) by retaining only a subset of visual tokens. Existing methods usually select tokens based on importance or redundancy. However, we observe that these criteria produce stable spatial biases across inputs and do not always outperform simple Uniform Grid sampling, highlighting the value of broad spatial coverage. Motivated by this, we propose S2Prune, a training-free pruning method that preserves spatial coverage while adapting token density to local image structure. We first divide the image into regions and assign at least one token to each region to preserve coverage. The remaining token budget is then distributed according to Laplacian variation, giving more tokens to regions with richer structure. We then use Early Representation Change (ERC), computed from the first decoder block, to select representative tokens within each region. We evaluate S2Prune across diverse settings and two MLLM architectures. On Qwen2.5-VL-7B-Instruct, it achieves the highest average accuracy among the evaluated training-free pruning methods. With only 32 of the original 576 visual tokens, it still retains 79.3% of the full-model performance. Code is available at https://github.com/yuanyuanjia71-spec/S2Prune.
Yuanyuan Jia, Shunpu Tang, Qianqian Yang
College of Information Science and Electronic Engineering, Zhejiang University Hangzhou, China