Organizations: Gaoling School of Artificial Intelligence, Renmin University of China · Baidu · Shanghai Jiao Tong University · Monash University · TierFlow Team · ELLIS Institute Tübingen · Max Planck Institute for Intelligent Systems · Tsinghua University
We introduce LoopVL to study whether Loop Transformers can be effectively extended to vision- language models. LoopVL combines Module-Loop and Model-Loop computation to iteratively update a unified vision-language state through shared modules. We train LoopVL from scratch through language pre-training, multimodal training, and post-training. LoopVL outperforms a range of similarly sized and larger non-recurrent models on multimodal understanding and visual reasoning benchmarks. We also observe Visual Aha Moments in LoopVL, characterized by pronounced shifts in visual attention across loops. LoopVL provides practical evidence for recurrent vision-language modeling and offers an intuitive perspective on how shared parameters can support deeper multimodal computation over continuously evolving visual-language states.
Figures & tables
Figure 2 : LoopVL architecture and the structure of its recurrent modules. (a) A visual encoder and projector supply language-space visual positions that enter the recurrent backbone alongside instruction tokens. Module-Loop repeats the complete L stack, and Model-Loop repeats the coupled L/H update cycle; the default schedule is [L × 3 → H] × 2. (b) Each L/H module contains 16 Transformer layers with normalization, gated self-attention, and a SwiGLU feed-forward path. L and H have separate parameters, each reused across its invocations.
Model
Recursions
FLOPs ( 1021 )
Tokens (T)
MMStar
RealWorld QA
VMCBench
AI2D
ChartQA
Math Vision
VisuLogic
MMK12
LoopVL
4 (H2L3)
2.47
0.14
63.47
70.98
70.90
75.49
74.52
38.49
27.00
49.65
Transformer-VL 1B
1
0.92
0.14
55.33
55.29
55.10
60.01
51.12
30.59
21.20
40.30
Transformer-VL 4B (Deep)
1
2.89
0.14
61.33
66.54
71.20
76.13
75.24
35.92
26.60
48.35
Transformer-VL 4B (Wide)
1
2.99
0.14
60.47
66.01
72.00
75.65
75.56
33.55
26.30
47.70
Table 1 : Comparison of LoopVL and Transformer-VL baselines at 0.14T training tokens. Bold denotes the highest score in each benchmark.
Figure 3 : Visual-state update controls. Normal is standard LoopVL training and inference. Block L visual updates and Freeze visual states are inference-only controls applied to the normal model: the former discards second-cycle L-module visual updates before they carry forward, whereas the latter keeps visual states fixed throughout the second cycle. Read-only trained is a separately trained variant that allows visual-state updates only in the first model cycle and uses the same read-only rule during training and inference. Visual tokens remain readable and nonvisual states continue updating.
Figure 4 : Our training-data sources and token allocations across the LoopVL training stages. HRM-Text Base denotes the open-source training data recipe from HRM-Text used for our language pretraining.
Figure 5 : Overview of our LoopVL training workflow. Dashed boxes denote models; solid boxes denote training steps. Language pretraining produces our own LLM checkpoint. Visual–language alignment and multimodal mid-training form LoopVL-Base, followed by supervised fine-tuning and visual reinforcement learning. Data sources and token allocations are shown in Figure 4 .
Configuration
Unrolled layers
MMStar
RealWorldQA
VMCBench
AI2D
ChartQA
H1L1
32
55.33
55.29
55.10
60.01
51.12
H1L3
64
58.13
59.61
60.50
66.06
60.40
H2L1
64
60.80
64.58
63.80
68.26
65.12
H2L3
128
63.47
70.98
70.90
75.49
74.52
Table 3 : Training-time H/L configuration ablation. Each variant is independently trained from initialization using its designated recurrence schedule throughout the full training pipeline. Unrolled layers denote the number of Transformer-layer calls in one forward pass.
Configuration
Unrolled layers
MMStar
RealWorldQA
VMCBench
AI2D
ChartQA
H1L1
32
0.47
0.00
0.50
1.62
0.00
H1L2
48
0.07
0.00
0.30
0.06
0.00
H1L3
64
0.07
0.00
0.10
0.45
0.04
H2L1
64
29.40
12.81
27.90
25.74
17.40
H2L2
96
51.33
58.95
63.10
65.38
63.28
H2L3
128
63.47
70.98
70.90
75.49
74.52
Table 4 : Inference-time H/L schedule sweep for the H2L3-trained model. H denotes the number of complete model cycles, and L denotes the number of L-module invocations within each cycle. Unrolled depth counts the total number of executed Transformer-layer calls in one forward pass. Darker and lighter shading indicate the best and second-best results, respectively.
Figure 6 : Spatial entropy over recurrent depth. The curve reports mean answer-side visual-attention entropy for the 32-sample diagnostic set. Lower values indicate a more concentrated spatial distribution. Entropy is generally lower during the second cycle, but fluctuates across module calls rather than decreasing monotonically. The horizontal axis uses zero-based executed-layer indices; loop markers identify repeated L calls and the cycle boundary.
Figure 7 : Timing of visual-attention concentration across the recurrent trajectory. The mean Gini curve uses instruction and reference-answer query positions in the 32-sample controlled diagnostic set. The marked model-cycle transition lies between zero-based indices 63 and 64, where mean Gini rises from 0.405 to 0.858. A high Gini value describes concentration; it does not by itself identify attention sinks or causally important visual tokens.
Figure 8 : Cross-cycle visual reallocation in eight selected examples. The first two rows compare the same final H-module layer at the two cycle endpoints (zero-based indices 63 and 127). The third marks tokens promoted from the first-cycle bottom half to the second-cycle top quartile: orange inside the outlined target region and cyan outside it. Each case uses a shared attention color scale across its two rounds. Reference answers identify the task, not intermediate model predictions.
Figure 9 : Visual-attention concentration under shared-module reuse. (a) Spatial entropy at the final layer of the first, second, and third L invocations (L1–L3), paired across model cycles. (b) Spatial entropy, 80% attention coverage, and Gini concentration at the final H layer. Gray lines show individual samples. Every blue mean uses all 32 samples. All four entropy panels share a common vertical scale. Entropy and coverage use reference-answer queries; Gini also includes instruction queries. See Appendix B for metric definitions and endpoint indices.
Figure 10 : Visual hidden-state dynamics. Left: Mean visual-token L2 updates across each complete 16-layer module invocation, with individual sample trajectories and their mean. Right: Cosine similarity between representations at the sampled module endpoints. Continued updates in the second cycle rule out an unchanged-copy description of these states.
Figure 11 : Logit-lens alignment for the first answer token. The analysis is based on a 32-sample state-probe collection. The curve shows the mean KL divergence from the final output distribution to each intermediate readout; the blue band represents the sample interquartile range, and pale orange marks H-module invocations. The final H call corresponds to the pronounced late approach of the model output distribution toward the endpoint. Depth counts use one-based indices.
Appendix figures & tables5 assets
Supplementary material from the paper’s appendix.
Appendix
Figure 12 : Additional visual reallocation examples, Set 1. Rows compare the same final H-module layer at the two cycle endpoints and mark tokens promoted from the first-cycle bottom half to the second-cycle top quartile. Orange and cyan mark promoted tokens inside and outside the outlined reference regions, respectively. Attention is normalized over the full visual grid separately in each cycle; paired heatmaps share a color scale. Images and overlays are jointly resized for display only. Labels give reference answers, not intermediate predictions.
Figure 13 : Additional visual reallocation examples, Set 2. Eight additional counting, size-comparison, and spatial-relation examples follow the same row layout, normalization, and promotion rule as Figure 12 . Rank promotion of individual tokens does not necessarily imply an increase in the total attention within the annotated target region.
Figure 14 : Additional visual reallocation examples, Set 3. Eight additional counting, size-comparison, and spatial-relation examples follow the same row layout, normalization, and promotion rule as Figure 12 . Rank promotion of individual tokens does not necessarily imply an increase in the total attention within the annotated target region.
Figure 15 : Additional visual reallocation examples, Set 4. Eight additional counting, size-comparison, and spatial-relation examples follow the same row layout, normalization, and promotion rule as Figure 12 . Rank promotion of individual tokens does not necessarily imply an increase in the total attention within the annotated target region.
Figure 16 : Additional visual reallocation examples, Set 5. Eight additional counting, size-comparison, and spatial-relation examples follow the same row layout, normalization, and promotion rule as Figure 12 . Rank promotion of individual tokens does not necessarily imply an increase in the total attention within the annotated target region.
Vision-Language Models (VLMs) frequently suffer from visual perception errors and hallucinations that compromise answer accuracy in complex reasoning tasks. Reinforcement Learning with Verifiable Rewards (RLVR) offers a promising solution by optimizing policies using answer correctness signals. Despite their effectiveness, prevailing RLVR methods face two critical limitations. First, much of the sampling budget is wasted on trajectories doomed to fail due to early visual description errors. Second, sparse rewards cannot distinguish whether failures stem from visual perception or reasoning stages. We introduce MIRL, a decoupled framework that addresses both limitations by leveraging mutual information (MI) between generated descriptions and visual inputs as a cheap pre-screening signal. This enables intelligent budget allocation toward high-potential trajectories via forking, while decoupled training provides independent MI-based rewards for visual perception optimization, resolving reward blindness. Experiments on six vision-language reasoning benchmarks demonstrate that MIRL achieves 70.22% average accuracy and successfully surpasses the performance of sampling 16 complete trajectories using only 10 pre-samples with top-6 selection (25% fewer complete trajectories). Our code is available at: https://anonymous.4open.science/r/mirl-main/.
Yin Zhang, Jiaxuan Zhao, Zonghan Wu +5
School of Mathematics, Tianjin University, Tianjin, China · Institute of Information Engineering, Chinese Academy of Sciences, Beijing, China · Shanghai Advanced Institute of Finance (SAIFS), East China Normal University, Shanghai, China +4
Goal-directed visual processing is a hallmark of human visual intelligence, resulting in representations that support downstream tasks such as categorization or search. Though vision-language models (VLMs) are often faced with these same tasks, their ability to recode visual representations when presented with goal-directed language remains poorly characterized. Indeed, prior work largely treats visual representations in VLMs as static repositories of visual information that are manipulated by language representations. In the present work, we provide evidence for two concrete instances of language-induced recoding of visual representations. First, we identify an abstract reference representation that denotes which objects are goal-relevant under a natural language prompt. We extract contrastive steering vectors corresponding to this reference representation and demonstrate that they are causally implicated in model predictions. These reference representations are abstract in that they generalize to different objects, different task contexts, and even from synthetic to naturalistic images. Second, we demonstrate language-induced attribute modulation: later layers selectively amplify goal-relevant attributes in visual representations of objects. We demonstrate this phenomenon across a range of different prompts. Finally, we provide a causal intervention that demonstrates that attribute modulation mediates a VLM's response distribution. Together, our results support a more dynamic account of cross-modality processing in VLMs -- rather than vision tokens serving as static repositories of information, they are modulated to support queries articulated in language.
Current Vision-Language-Action (VLA) models typically treat the deepest representation of a vision-language backbone as universally optimal for action prediction. However, robotic manipulation is composed of many frequent closed-loop spatial adjustments, for which excessive abstraction may waste computation and weaken low-level geometric cues essential for precise control. Existing early-exit strategies attempt to reduce computation by stopping at predefined layers or applying heuristic rules such as action consistency, but they do not directly answer when a representation is actually sufficient for action. In this paper, we present LoopVLA, a recurrent VLA architecture that jointly learns representation refinement, action prediction, and sufficiency estimation. LoopVLA iteratively applies a shared Transformer block to refine multimodal tokens, and at each iteration produces both a candidate action and a sufficiency score that estimates whether further refinement is necessary. By sharing parameters across iterations, LoopVLA decouples refinement from absolute layer indices and grounds sufficiency estimation in the evolving representation itself. Since sufficiency has no direct supervision, we introduce a self-supervised distribution alignment objective, where intermediate confidence scores are trained to match the relative action quality across refinement steps, thereby linking sufficiency learning to policy optimization signals. Experiments on LIBERO, LIBERO-Plus, and VLA-Arena show that LoopVLA pushes the efficiency-performance frontier of VLA policies, reducing parameters by 45% and improving inference throughput by up to 1.7 times while matching or outperforming strong baselines in task success.
Boyang Shen, Kaixiang Yang, Hao Wang +4
1Huazhong University of Science and Technology · 2Wuhan United Imaging Surgical Co.,Ltd. (UIS)