cs.CVOct 6, 2026

VisionWeave: Weaving Elastic Visual Representations as a Native Capability of MLLMs

Authors: Yuan Feng, Qize Yang, Ruizhe Chen, Sibo Song, Haolin He, Muzhi Zhu, Zihan Liu, Yunfei Chu, +4 more

Organizations: University of Science and Technology of China · Alibaba Token Hub, Alibaba Group · The Chinese University of Hong Kong · Zhejiang University

Abstract

Multimodal large language models have become the dominant paradigm for visual understanding, but incur substantial costs by encoding inputs into dense, fixed-size patch tokens. However, visual information is unevenly distributed: some regions require fine-grained detail, while others admit compact representations. Downsampling sacrifices this detail, while existing token pruning and adaptive approaches remain limited in content-adaptive granularity, task generalization, and integration with modern MLLMs and serving infrastructure. Overcoming these limitations calls for foundation models that learn, end to end, where-and at what granularity-to allocate visual representations, a native capability we term elastic visual representation weaving. We introduce VisionWeave, establishing this capability in frontier-level MLLMs through large-scale training. It combines two components: a gated spatial pooler constructs coarse-grained representations alongside native fine-grained representations within a shared MRoPE coordinate, while a granularity router learns their content-adaptive allocation. Through self-distillation alone, we validate this capability on Qwen3.5-4B and scale to Qwen3.8-27B with over 30K A100 GPU-hours. Based on Qwen3.8-27B, VisionWeave adaptively adjusts token savings to visual content, saving 43.0% tokens on average while retaining 98.9% native performance across eight benchmarks, versus only 88% performance preserved for token pruning baselines with a fixed 50% savings target. Extensive evaluations confirm robust efficiency-quality trade-offs across diverse tasks, resolutions and video frames. When deployed on SGLang serving engine, our method achieves a 2.3x throughput gain while reducing mean TTFT by 54.4% and mean TPOT by 60.6%. Together, we believe these results position elastic visual weaving as a promising capability for next-generation multimodal models.

Figures & tables

Appendix figures & tables10 assets

Supplementary material from the paper’s appendix.

Appendix

Explore similar work

CardsList
  1. Just MLPs: Efficient Visual State Reconstruction for Multimodal Language Models

    Sep 28, 2026Jingdi lei, Junxian Li, Di Zhang +3Long Visual-Token SequencesMultimodal Large Language Models

  2. SepPrune:A Separator-based Pruning Framework for Efficient Multimodal Large Language Models

    Jul 28, 2026Yuchen Wang, Qihui Zhu, Yang Liu +2Visual Token PruningLong Visual-Token Sequences

  3. S2^2Prune: Spatially Structured Visual Token Pruning for Multimodal Large Language Models

    Sep 1, 2026Yuanyuan Jia, Shunpu Tang, Qianqian YangVisual Token PruningMultimodal Large Language Models