cs.CVSep 29, 2026

LoopVL: Recurrent Visual Intelligence

Authors: Zhe Qian, Ziyang Gong, Zhongxing Xu, Hehan Li, Zhonghua Wang, Fei Luo, Mingxuan Wang, Xue Yang, +4 more

Organizations: Gaoling School of Artificial Intelligence, Renmin University of China · Baidu · Shanghai Jiao Tong University · Monash University · TierFlow Team · ELLIS Institute Tübingen · Max Planck Institute for Intelligent Systems · Tsinghua University

Abstract

We introduce LoopVL to study whether Loop Transformers can be effectively extended to vision- language models. LoopVL combines Module-Loop and Model-Loop computation to iteratively update a unified vision-language state through shared modules. We train LoopVL from scratch through language pre-training, multimodal training, and post-training. LoopVL outperforms a range of similarly sized and larger non-recurrent models on multimodal understanding and visual reasoning benchmarks. We also observe Visual Aha Moments in LoopVL, characterized by pronounced shifts in visual attention across loops. LoopVL provides practical evidence for recurrent vision-language modeling and offers an intuitive perspective on how shared parameters can support deeper multimodal computation over continuously evolving visual-language states.

Figures & tables

Appendix figures & tables5 assets

Supplementary material from the paper’s appendix.

Appendix

Explore similar work

CardsList
  1. MIRL: Mutual Information-Guided Reinforcement Learning for Vision-Language Models

    May 2, 2026Yin Zhang, Jiaxuan Zhao, Zonghan Wu +5Recent Vision-Language ModelsReinforcement Learning With Verifiable Reward

  2. Linguistic Context Recodes Visual Representations in Vision-Language Models

    Jul 21, 2026Brian Song, Michael A. Lepori, Ellie PavlickVisual RepresentationsCross-Modality

  3. LoopVLA: Learning Sufficiency in Recurrent Refinement for Vision-Language-Action Models

    May 11, 2026Boyang Shen, Kaixiang Yang, Hao Wang +4Action PredictionPhotometric Supervision