Organizations: University of Science and Technology of China · Alibaba Token Hub, Alibaba Group · The Chinese University of Hong Kong · Zhejiang University
Multimodal large language models have become the dominant paradigm for visual understanding, but incur substantial costs by encoding inputs into dense, fixed-size patch tokens. However, visual information is unevenly distributed: some regions require fine-grained detail, while others admit compact representations. Downsampling sacrifices this detail, while existing token pruning and adaptive approaches remain limited in content-adaptive granularity, task generalization, and integration with modern MLLMs and serving infrastructure. Overcoming these limitations calls for foundation models that learn, end to end, where-and at what granularity-to allocate visual representations, a native capability we term elastic visual representation weaving. We introduce VisionWeave, establishing this capability in frontier-level MLLMs through large-scale training. It combines two components: a gated spatial pooler constructs coarse-grained representations alongside native fine-grained representations within a shared MRoPE coordinate, while a granularity router learns their content-adaptive allocation. Through self-distillation alone, we validate this capability on Qwen3.5-4B and scale to Qwen3.8-27B with over 30K A100 GPU-hours. Based on Qwen3.8-27B, VisionWeave adaptively adjusts token savings to visual content, saving 43.0% tokens on average while retaining 98.9% native performance across eight benchmarks, versus only 88% performance preserved for token pruning baselines with a fixed 50% savings target. Extensive evaluations confirm robust efficiency-quality trade-offs across diverse tasks, resolutions and video frames. When deployed on SGLang serving engine, our method achieves a 2.3x throughput gain while reducing mean TTFT by 54.4% and mean TPOT by 60.6%. Together, we believe these results position elastic visual weaving as a promising capability for next-generation multimodal models.
Figures & tables
Figure 1: Weaving elastic visual representations. More examples appear in Appendix C .
Figure 2: VisionWeave achieves content-adaptive token savings with near-native quality, whereas fix-budget baselines degrade sharply on challenging tasks. Visual case studies on ScreenSpotV2 and DocVQA illustrate this difference (Appendix B ). Our VisionWeave still retains higher quality even when baselines are granted its adaptive saving ratios (Appendix E.3 ).
Figure 3: Overall architecture of VisionWeave. The gated spatial pooler complements native fine-grained tokens Vfine with coarse-grained representations Vcoarse within the same MRoPE coordinate, while the granularity router adaptively weaves them into a mixed-granularity sequence for the LLM.
Stage 1
Stage 2
Stage 3
Trainable components
Gated Spatial Pooler
Granularity Router
Pooler, LLM
Teacher
Native model
Native model
Native model
Routing
Bypassed (all coarse)
Soft-mixing routing
Hard routing
Loss terms
Ldistill
Ldistill+0.02Lbal
Ldistill
Student visual tokens
25%
100%
Content-adaptive
Table 1: Three-stage self-distillation for elastic visual processing.
Source
Domain
Approx. training samples
LLaVA-OneVision ( An et al., 2025 )
Image understanding
246K
LLaVA-Video-178K ( Zhang et al., 2025b )
Video understanding
281K
UGround ( Gou et al., 2025 )
GUI grounding
250K
Total
777K
Table 2: Training-data composition for Qwen3.8-27B.
Setting
Stage 1
Stage 2
Stage 3
Peak learning rate
1×10−4
1×10−5
3×10−6
Minimum learning rate
5×10−6
0
0
Warmup fraction
5%
5%
2%
Global batch size
32
64
512
Training steps
5,257
2,500
1,000
Gradient clipping
1.0
1.0
0.5
Table 3: Training hyperparameters for Qwen3.8-27B.
Figure 4: Training dynamics of three-stage self-distillation.
Native Model
Fixed Savings (s=50%)
Content-adaptive Savings
Qwen3.5-4B
512 token/Image
Downsampling
FastV †
VisionZip
VisionWeave
Savings (%)
RealWorldQA
74.64
72.16 ( ↓ 3.32%)
70.98 ( ↓ 4.90%)
70.72 ( ↓ 5.25%)
73.46 ( ↓ 1.58%)
43.2
DocVQA
91.92
82.71 ( ↓ 10.02%)
78.69 ( ↓ 14.39%)
76.57 ( ↓ 16.70%)
91.03 ( ↓ 0.97%)
16.0
InfoVQA
67.44
52.87 ( ↓ 21.60%)
51.75 ( ↓ 23.27%)
50.89 ( ↓ 24.54%)
66.11 ( ↓ 1.97%)
10.6
ScreenSpotV2
89.86
80.42 ( ↓ 10.51%)
72.56 ( ↓ 19.25%)
70.44 ( ↓ 21.61%)
89.47 ( ↓ 0.43%)
24.1
HallusionBench
69.88
69.09 ( ↓ 1.13%)
69.26 ( ↓ 0.89%)
68.64 ( ↓ 1.77%)
69.26 ( ↓ 0.89%)
44.0
Table 4: Content-adaptive token savings with near-native quality.
Figure 5: VisionWeave achieves consistently favorable trade-offs across tasks and input resolutions.
Figure 10
Model
Throughput (req/min) ↑
Mean TTFT (s) ↓
P95 TTFT (s) ↓
Mean TPOT (ms) ↓
P95 TPOT (ms) ↓
Native Model
0.90
169.90
281.41
355.64
715.58
VisionWeave
2.07 ( 2.30× )
77.54 ( ↓ 54.4%)
120.59 ( ↓ 57.1%)
140.29 ( ↓ 60.6%)
289.19 ( ↓ 59.6%)
Table 5: VisionWeave accelerates Qwen3.8-27B serving on SGLang.
Figure 8: VisionWeave enables flexible inference-time control of token savings while closely matching native-model performance when savings are disabled.
Qwen3.5-4B
Qwen3.8-27B
Benchmark
VisionWeave
Shuffled
Δ
VisionWeave
Shuffled
Δ
RealWorldQA
73.46
73.46
0.00
77.91
76.60
−1.31
DocVQA
91.03
88.70
−2.32
91.33
88.35
−2.98
InfoVQA
66.11
64.83
−1.28
71.61
69.04
−2.58
ScreenSpotV2
89.47
80.03
−9.43
91.43
59.28
−32.15
HallusionBench
69.26
67.32
−1.95
73.60
71.66
−1.95
Table 6: Learned granularity allocation improves performance over shuffled routing.
MoE Key Lab of BIPC, NEL-BITA, University of Science and Technology of China, Hefei, China · ChangXin Memory Technologies, Hefei, China · APKL of BIIP, IAI, Hefei Comprehensive National Science Center, Hefei, China