Active multimodal agents use visual tools to acquire task-relevant evidence while reasoning. Although reinforcement learning samples multiple interaction trajectories per input, outcome-based objectives primarily use the group to estimate scalar advantages, leaving complementary visual discoveries underused. We introduce VISTA, which internalizes collective visual experience through on-policy distillation by turning observations from same-input rollouts into shared supervision. Collective visual experience distillation (CVED) organizes these observations with their interaction context and aligns them with individual decisions, while heterogeneity-aware policy improvement (HAPI) reinforces successful trajectories and provides experience-guided distillation for unsuccessful attempts. An experience-conditioned teacher evaluates the student's sampled response prefixes, allowing discoveries from one trajectory to guide learning in another without replacing the student's original history or generating new target trajectories. The trained agent retains its visual tools and acts using its own interaction history. VISTA achieves the strongest average performance among the evaluated active multimodal agents of comparable size and consistently outperforms same-backbone training baselines across fine-grained perception and general reasoning tasks, demonstrating the value of collective experience for active multimodal learning.
Figures & tables
Figure 1: VISTA combines experience transfer with heterogeneous policy improvement. The complete rollout group supplies experience for CVED , while HAPI allocates reinforcement and distillation according to trajectory outcomes. The teacher scores the student’s sampled responses without generating new training trajectories.
Model
V* Bench
ZoomBench
HR-Bench-4K
HR-Bench-8K
MME-RW-EN
MME-RW-CN
Average
Closed-source models
GPT-5.2
79.06
50.89
81.12
78.38
72.60
68.80
71.81
GPT-5.4
76.96
52.66
84.00
77.88
74.20
70.93
72.77
Gemini-2.5-Pro
78.01
49.82
80.63
77.50
71.29
69.11
71.06
Gemini-3-Flash
78.53
50.06
80.87
78.00
71.35
69.38
71.37
Open-source models
Table 1: Performance comparison on visual perception and reasoning benchmarks. We report accuracy (%). Δ denotes the gain over the corresponding Qwen3-VL base model.
Methods
Fine-Grained Visual Tasks
General Reasoning Tasks
Average
V* Bench
HR-Bench-4K
HR-Bench-8K
MMStar
MMBench
MathVista
Qwen3-VL-4B
Base
81.68
78.50
76.25
68.73
81.22
73.70
76.68
GRPO
84.12
80.08
77.50
70.87
82.48
75.10
78.36
OPSD
85.34
81.25
78.25
71.53
82.91
75.60
79.15
VCSD
83.77
80.38
76.88
69.20
83.08
74.55
77.98
Table 2: Comparison of training methods on fine-grained visual tasks and general reasoning tasks. We report accuracy (%) for Qwen3-VL-4B and Qwen3-VL-8B.
Setting
V*
HR-4K
HR-8K
Avg.
Answer-only
86.56
81.30
77.80
81.89
Unstructured
86.91
82.13
78.25
82.43
Oracle ROI
87.96
82.80
79.00
83.25
CVED
88.48
83.50
79.50
83.83
Table 3: CVED ablation: privileged supervision under fixed HAPI routing.
Setting
V*
HR-4K
HR-8K
Avg.
Answer-only
86.56
81.30
77.80
81.89
Unstructured
86.91
82.13
78.25
82.43
Oracle ROI
87.96
82.80
79.00
83.25
CVED
88.48
83.50
79.50
83.83
Table 3: CVED ablation: privileged supervision under fixed HAPI routing.
Setting
V*
HR-4K
HR-8K
Avg.
GRPO-only
84.12
80.08
77.50
80.57
CVED-only
86.21
81.50
78.13
81.95
Uniform hybrid
87.43
82.38
78.80
82.87
HAPI
88.48
83.50
79.50
83.83
Table 4: HAPI ablation: learning rules under fixed experience construction.
Teacher
V*Bench
ZoomBench
HR-Bench-4K
HR-Bench-8K
MME-RW-EN
MME-RW-CN
Avg.
Frozen initial
87.43
51.95
83.38
80.75
68.19
69.51
73.54
EMA (ours)
88.48
53.85
83.50
79.50
68.24
69.29
73.81
Table 5: Teacher-update ablation on Qwen3-VL-4B.
Objective
V*Bench
ZoomBench
HR-Bench-4K
HR-Bench-8K
MME-RW-EN
MME-RW-CN
Avg.
Forward KL
86.91
50.89
82.25
77.80
67.10
68.10
72.18
Reverse KL
87.96
53.96
82.95
79.75
67.80
69.05
73.58
JSD
88.48
53.85
83.50
79.50
68.24
69.29
73.81
Table 6: Divergence-objective comparison on Qwen3-VL-4B with fixed CVED and HAPI.
Figure 2: Sensitivity of VISTA on Qwen3-VL-4B to (a) rollout-group size K , (b) distillation weight λ , and (c) visual artifact budget B . Points report V*Bench, HR-4K, and HR-8K accuracy. Tested settings are equally spaced along each horizontal axis.
Appendix figures & tables7 assets
Supplementary material from the paper’s appendix.
Appendix
Channel
Content
Main setting
Visual observations
selected images from the rollout group
enabled
Interaction context
producing decisions and artifact relations
enabled
Outcome feedback
trajectory status and failure category
enabled
Appendix
Table 8: Additional teacher context used by CVED. Generated final answers and reference answers are not included.
Figure 3: Student system prompt used for active visual rollouts.
Figure 4: Collective-experience teacher instruction. The teacher evaluates the student’s unchanged response under training-only visual hindsight.
Figure 5: Teacher-context serialization with visual experience, interaction context, and outcome feedback. Braced variables are instantiated for each input.
Parameter
Value
Optimization and training duration
Optimizer
AdamW
Learning rate
1×10−6
Weight decay
0.01
Optimizer steps
65
Equivalent data passes (epochs)
≈1
Appendix
Table 9: Default VISTA training configuration for both backbone sizes. Equivalent data passes are computed as optimizer steps times input questions per batch divided by training-set size.
Experience source
V*
HR-4K
HR-8K
Avg.
Self-only hindsight
85.34
82.25
77.50
81.70
Collective visual experience
88.48
83.50
79.50
83.83
Appendix
Table 10: Experience-source ablation on Qwen3-VL-4B. Accuracy (%) is averaged over three seeds; Avg. is the unweighted mean across the three benchmarks. The collective-experience row reproduces the default CVED results in Table 4 . Bold and underlined values mark the highest and lowest scores in each column, respectively.
Setting
RL subset
Distillation subset
V*
HR-4K
HR-8K
Avg.
Uniform hybrid
All
All
87.43
82.38
78.80
82.87
RL mask only
Successful
All
86.91
82.50
78.25
82.55
Distillation mask only
All
Unsuccessful
87.43
82.25
78.92
82.87
HAPI
Successful
Unsuccessful
88.48
83.50
79.50
83.83
Appendix
Table 11: Crossed routing ablation on Qwen3-VL-4B. We report three-seed mean accuracy (%) and the unweighted benchmark average. Uniform hybrid and HAPI reproduce Table 4 . Bold and underlined values mark the highest and lowest scores in each column, respectively.
We show that multimodal models possess strong reasoning abilities and that an appropriate harness can unlock their potential to solve tasks across diverse interactive environments. We introduce VISTA, a visual harness that gives a general-purpose multimodal model long-horizon vision. VISTA allows the model to directly perceive the environment through visual observations and maintains a lossless visual memory that preserves past observations in their original form. The model can actively retrieve these observations and reorganize its visual input as it reasons. On ARC-AGI-3, VISTA improves Claude Opus 5.0's Relative Human Action Efficiency score from 40.68 to a perfect 100.00, with the model completing all 25 public games using 57.4% fewer actions than first-time human participants. VISTA's simple design also allows it to extend naturally to diverse visual environments with minimal adaptation. Across three additional benchmarks covering a diverse range of visual games and puzzles, it substantially outperforms baselines using the same underlying model with minimal harnesses. Our results highlight VISTA's potential as a general-purpose visual harness for advancing multimodal agents in complex visual environments.
Visual agents solve problems by interleaving reasoning with image operations, and on-policy distillation (OPD) provides guidance from a strong teacher on student-generated interaction trajectories. However, image operations change the evidence available for subsequent reasoning, so local errors in evidence acquisition (Acquire), reading (Read), or answer grounding (Ground) can propagate through the trajectory and lead to incorrect answers. Existing multimodal OPD methods primarily construct or contrast auxiliary views of the original image to strengthen supervision, without explicitly modeling the connections between student actions, resulting observations, and subsequent reasoning. This limits their ability to provide corrections tailored to different failure stages. We introduce Reflection on Visual Evidence (ReVuE), an on-policy distillation method for visual agents. ReVuE compares multiple student-generated trajectories for the same query, summarizes the observed visual evidence, and diagnoses the first failure across the Acquire, Read, and Ground stages. The resulting reflections provide training-time context for the teacher. We group and reweight token-level distillation losses according to how strongly these reflections affect the teacher's predictions. This design translates trajectory-level evidence diagnosis into targeted token-level supervision, guiding students to improve their visual evidence acquisition and reasoning. Across 11 benchmarks spanning the Qwen2.5-VL and InternVL3.5 model families, ReVuE outperforms all evaluated OPD baselines in weighted-average scores for perception, mathematical reasoning, and general tasks. ReVuE also reduces redundancy in reasoning and tool calls while improving tool-call accuracy and task accuracy. Code is available at https://github.com/sylvain-wei/ReVuE
Multimodal agents can now tackle complex reasoning tasks with diverse tools, yet they still suffer from inefficient tool use and inflexible orchestration in open-ended settings. A central challenge is enabling such agents to continually improve without parameter updates by learning from past trajectories. We identify two complementary forms of reusable knowledge essential for this goal: experiences, providing concise action-level guidance for tool selection and decision making, and skills, providing structured task-level guidance for planning and tool use. To this end, we propose XSkill, a dual-stream framework for continual learning from experience and skills in multimodal agents. XSkill grounds both knowledge extraction and retrieval in visual observations. During accumulation, XSkill distills and consolidates experiences and skills from multi-path rollouts via visually grounded summarization and cross-rollout critique. During inference, it retrieves and adapts this knowledge to the current visual context and feeds usage history back into accumulation to form a continual learning loop. Evaluated on five benchmarks across diverse domains with four backbone models, XSkill consistently and substantially outperforms both tool-only and learning-based baselines. Further analysis reveals that the two knowledge streams play complementary roles in influencing the reasoning behaviors of agents and show superior zero-shot generalization.
Guanyu Jiang, Zhaochen Su, Xiaoye Qu +1
Hong Kong University of Science and Technology · Zhejiang University · Huazhong University of Science and Technology