cs.CVSep 28, 2026

Just MLPs: Efficient Visual State Reconstruction for Multimodal Language Models

Authors: Jingdi lei, Junxian Li, Di Zhang, Zhanqiu Zhang, Yiwen Guo, Soujanya Poria

Organizations: Nanyang Technological University · Shanghai Jiao Tong University · Fudan University · LIGHTSPEED · Independent Researcher

Abstract

Long visual token sequences often account for a substantial fraction of the computational overhead in multimodal large language models~(MLLMs). Existing approaches reduce this cost by pruning redundant visual tokens, but permanently discard visual evidence that may become useful in subsequent layers. We instead ask whether all visual tokens can be preserved while reducing the cost of repeatedly evolving the representations through the Transformer. To answer this question, we perform low-rank interventions on visual-to-text information flow. We find that, after visual-to-text attention is blocked, restoring only a few directions recovers most of the lost accuracy, suggesting the relevant visual influence is concentrated in a low-dimensional subspace. We further observe strong predictability in layer-specific visual states: lightweight MLPs approximate them with high cosine similarity and low reconstruction error. Motivated by these findings, we propose δδ-Vision, which replaces repeated Transformer evolution of visual tokens with lightweight low-rank adapters that construct layer-wise visual memories while preserving all visual tokens for text retrieval. Across image and video benchmarks, δδ-Vision achieves higher accuracy than visual token pruning baselines at comparable or lower computation, while delivering competitive inference efficiency without discarding visual tokens.

Figures & tables

Appendix figures & tables8 assets

Supplementary material from the paper’s appendix.

Appendix

Explore similar work

CardsList
  1. Learning to Predict Middle-Layer Attention in MLLMs for Visual Token Pruning

    Aug 4, 2026Yuyao Sun, Tao Deng, Shuang Li +3Visual Token PruningRecent Vision-Language Models

  2. Attend, Transform, or Silence: Operator-Level Visual Skipping for Efficient Multimodal LLM Inference

    Jun 30, 2026Zhaoyang Luo, Runmin Dong, Miao Yang +4Long Visual-Token SequencesMultimodal Large Language Models

  3. A More Word-like Image Tokenization for MLLMs

    May 18, 2026Hyun Lee, Hyemin Jeong, Yejin Kim +4Long Visual-Token SequencesVisual Tokenizers