Organizations: University of Science and Technology of China · Alibaba Token Hub, Alibaba Group · The Chinese University of Hong Kong · Zhejiang University
Multimodal large language models have become the dominant paradigm for visual understanding, but incur substantial costs by encoding inputs into dense, fixed-size patch tokens. However, visual information is unevenly distributed: some regions require fine-grained detail, while others admit compact representations. Downsampling sacrifices this detail, while existing token pruning and adaptive approaches remain limited in content-adaptive granularity, task generalization, and integration with modern MLLMs and serving infrastructure. Overcoming these limitations calls for foundation models that learn, end to end, where-and at what granularity-to allocate visual representations, a native capability we term elastic visual representation weaving. We introduce VisionWeave, establishing this capability in frontier-level MLLMs through large-scale training. It combines two components: a gated spatial pooler constructs coarse-grained representations alongside native fine-grained representations within a shared MRoPE coordinate, while a granularity router learns their content-adaptive allocation. Through self-distillation alone, we validate this capability on Qwen3.5-4B and scale to Qwen3.8-27B with over 30K A100 GPU-hours. Based on Qwen3.8-27B, VisionWeave adaptively adjusts token savings to visual content, saving 43.0% tokens on average while retaining 98.9% native performance across eight benchmarks, versus only 88% performance preserved for token pruning baselines with a fixed 50% savings target. Extensive evaluations confirm robust efficiency-quality trade-offs across diverse tasks, resolutions and video frames. When deployed on SGLang serving engine, our method achieves a 2.3x throughput gain while reducing mean TTFT by 54.4% and mean TPOT by 60.6%. Together, we believe these results position elastic visual weaving as a promising capability for next-generation multimodal models.
Figures & tables
Figure 1: Weaving elastic visual representations. More examples appear in Appendix C .
Figure 2: VisionWeave achieves content-adaptive token savings with near-native quality, whereas fix-budget baselines degrade sharply on challenging tasks. Visual case studies on ScreenSpotV2 and DocVQA illustrate this difference (Appendix B ). Our VisionWeave still retains higher quality even when baselines are granted its adaptive saving ratios (Appendix E.3 ).
Figure 3: Overall architecture of VisionWeave. The gated spatial pooler complements native fine-grained tokens Vfine with coarse-grained representations Vcoarse within the same MRoPE coordinate, while the granularity router adaptively weaves them into a mixed-granularity sequence for the LLM.
Stage 1
Stage 2
Stage 3
Trainable components
Gated Spatial Pooler
Granularity Router
Pooler, LLM
Teacher
Native model
Native model
Native model
Routing
Bypassed (all coarse)
Soft-mixing routing
Hard routing
Loss terms
Ldistill
Ldistill+0.02Lbal
Ldistill
Student visual tokens
25%
100%
Content-adaptive
Table 1: Three-stage self-distillation for elastic visual processing.
Source
Domain
Approx. training samples
LLaVA-OneVision ( An et al., 2025 )
Image understanding
246K
LLaVA-Video-178K ( Zhang et al., 2025b )
Video understanding
281K
UGround ( Gou et al., 2025 )
GUI grounding
250K
Total
777K
Table 2: Training-data composition for Qwen3.8-27B.
Setting
Stage 1
Stage 2
Stage 3
Peak learning rate
1×10−4
1×10−5
3×10−6
Minimum learning rate
5×10−6
0
0
Warmup fraction
5%
5%
2%
Global batch size
32
64
512
Training steps
5,257
2,500
1,000
Gradient clipping
1.0
1.0
0.5
Table 3: Training hyperparameters for Qwen3.8-27B.
Figure 4: Training dynamics of three-stage self-distillation.
Native Model
Fixed Savings (s=50%)
Content-adaptive Savings
Qwen3.5-4B
512 token/Image
Downsampling
FastV †
VisionZip
VisionWeave
Savings (%)
RealWorldQA
74.64
72.16 ( ↓ 3.32%)
70.98 ( ↓ 4.90%)
70.72 ( ↓ 5.25%)
73.46 ( ↓ 1.58%)
43.2
DocVQA
91.92
82.71 ( ↓ 10.02%)
78.69 ( ↓ 14.39%)
76.57 ( ↓ 16.70%)
91.03 ( ↓ 0.97%)
16.0
InfoVQA
67.44
52.87 ( ↓ 21.60%)
51.75 ( ↓ 23.27%)
50.89 ( ↓ 24.54%)
66.11 ( ↓ 1.97%)
10.6
ScreenSpotV2
89.86
80.42 ( ↓ 10.51%)
72.56 ( ↓ 19.25%)
70.44 ( ↓ 21.61%)
89.47 ( ↓ 0.43%)
24.1
HallusionBench
69.88
69.09 ( ↓ 1.13%)
69.26 ( ↓ 0.89%)
68.64 ( ↓ 1.77%)
69.26 ( ↓ 0.89%)
44.0
Table 4: Content-adaptive token savings with near-native quality.
Figure 5: VisionWeave achieves consistently favorable trade-offs across tasks and input resolutions.
Figure 10
Model
Throughput (req/min) ↑
Mean TTFT (s) ↓
P95 TTFT (s) ↓
Mean TPOT (ms) ↓
P95 TPOT (ms) ↓
Native Model
0.90
169.90
281.41
355.64
715.58
VisionWeave
2.07 ( 2.30× )
77.54 ( ↓ 54.4%)
120.59 ( ↓ 57.1%)
140.29 ( ↓ 60.6%)
289.19 ( ↓ 59.6%)
Table 5: VisionWeave accelerates Qwen3.8-27B serving on SGLang.
Figure 8: VisionWeave enables flexible inference-time control of token savings while closely matching native-model performance when savings are disabled.
Qwen3.5-4B
Qwen3.8-27B
Benchmark
VisionWeave
Shuffled
Δ
VisionWeave
Shuffled
Δ
RealWorldQA
73.46
73.46
0.00
77.91
76.60
−1.31
DocVQA
91.03
88.70
−2.32
91.33
88.35
−2.98
InfoVQA
66.11
64.83
−1.28
71.61
69.04
−2.58
ScreenSpotV2
89.47
80.03
−9.43
91.43
59.28
−32.15
HallusionBench
69.26
67.32
−1.95
73.60
71.66
−1.95
Table 6: Learned granularity allocation improves performance over shuffled routing.
Long visual token sequences often account for a substantial fraction of the computational overhead in multimodal large language models~(MLLMs). Existing approaches reduce this cost by pruning redundant visual tokens, but permanently discard visual evidence that may become useful in subsequent layers. We instead ask whether all visual tokens can be preserved while reducing the cost of repeatedly evolving the representations through the Transformer. To answer this question, we perform low-rank interventions on visual-to-text information flow. We find that, after visual-to-text attention is blocked, restoring only a few directions recovers most of the lost accuracy, suggesting the relevant visual influence is concentrated in a low-dimensional subspace. We further observe strong predictability in layer-specific visual states: lightweight MLPs approximate them with high cosine similarity and low reconstruction error. Motivated by these findings, we propose δ-Vision, which replaces repeated Transformer evolution of visual tokens with lightweight low-rank adapters that construct layer-wise visual memories while preserving all visual tokens for text retrieval. Across image and video benchmarks, δ-Vision achieves higher accuracy than visual token pruning baselines at comparable or lower computation, while delivering competitive inference efficiency without discarding visual tokens.
Jingdi lei, Junxian Li, Di Zhang +3
Nanyang Technological University · Shanghai Jiao Tong University · Fudan University +2
Processing long visual token sequences from high-resolution images makes multi-step reasoning computationally expensive for multimodal Large Language Models (MLLMs). Existing one-shot pruning and aggregation methods compress visual tokens into a fixed context before decoding. However, visual evidence needs can shift as reasoning unfolds, making it difficult for a fixed compressed context to retain all the details needed across stages. To address this challenge, we propose ViMoD, a lightweight framework that maintains a compact visual context while preserving access to original fine-grained evidence as reasoning needs evolve. Deformable Aggregation of Region-wise Tokens (DART) learns content-adaptive groups and aggregation capacities, constructing compact Coarse representations linked to recoverable original Fine tokens. Temporal Routing for Adaptive Contextual Evidence (TRACE) integrates decoding history to anticipate upcoming evidence needs and select, retain, or replace active Fine-token groups. Selected Fine tokens augment the persistent Coarse context in the frozen backbone, enabling stage-specific evidence access without continuously attending to all visual tokens. On Qwen3-VL-4B, ViMoD outperforms all evaluated baselines on all eight reasoning benchmarks at a 20% target visual token budget, improving the mean normalized score by 39.0% over the strongest evaluated one-shot baseline. These gains are achieved with only 0.0546% additional trainable parameters relative to the frozen backbone.
Yicheng Xue, Han Wu, Jufeng Yang +4
City University of Hong Kong · Zhejiang University · Peking University +2
Recent multimodal large language models (MLLMs), such as Qwen2.5-VL and InternVL3, generate large numbers of vision tokens for high-resolution inputs, leading to substantial computational cost. Existing vision token pruning methods either depend on cross-modal attention and cannot prune before the prefill stage, or rely on diversity estimation with high computational overhead. We observe that attention scores from both vision and text tokens peak at modality separator tokens, suggesting that these separators bridge the two modalities. Based on this observation, we propose SepPrune, an efficient, training-free, plug-and-play pruning method that uses the separator token as a unified query to rank and select informative vision tokens. SepPrune reuses the LLM's built-in projection parameters and requires no architectural changes. Experiments on Qwen2.5-VL-7B show that SepPrune achieves state-of-the-art performance, retaining 96.3% of the original accuracy while removing 80.2% of vision tokens.
Yuchen Wang, Qihui Zhu, Yang Liu +2
MoE Key Lab of BIPC, NEL-BITA, University of Science and Technology of China, Hefei, China · ChangXin Memory Technologies, Hefei, China · APKL of BIIP, IAI, Hefei Comprehensive National Science Center, Hefei, China