Active multimodal agents use visual tools to acquire task-relevant evidence while reasoning. Although reinforcement learning samples multiple interaction trajectories per input, outcome-based objectives primarily use the group to estimate scalar advantages, leaving complementary visual discoveries underused. We introduce VISTA, which internalizes collective visual experience through on-policy distillation by turning observations from same-input rollouts into shared supervision. Collective visual experience distillation (CVED) organizes these observations with their interaction context and aligns them with individual decisions, while heterogeneity-aware policy improvement (HAPI) reinforces successful trajectories and provides experience-guided distillation for unsuccessful attempts. An experience-conditioned teacher evaluates the student's sampled response prefixes, allowing discoveries from one trajectory to guide learning in another without replacing the student's original history or generating new target trajectories. The trained agent retains its visual tools and acts using its own interaction history. VISTA achieves the strongest average performance among the evaluated active multimodal agents of comparable size and consistently outperforms same-backbone training baselines across fine-grained perception and general reasoning tasks, demonstrating the value of collective experience for active multimodal learning.
Figures & tables
Figure 1: VISTA combines experience transfer with heterogeneous policy improvement. The complete rollout group supplies experience for CVED , while HAPI allocates reinforcement and distillation according to trajectory outcomes. The teacher scores the student’s sampled responses without generating new training trajectories.
Model
V* Bench
ZoomBench
HR-Bench-4K
HR-Bench-8K
MME-RW-EN
MME-RW-CN
Average
Closed-source models
GPT-5.2
79.06
50.89
81.12
78.38
72.60
68.80
71.81
GPT-5.4
76.96
52.66
84.00
77.88
74.20
70.93
72.77
Gemini-2.5-Pro
78.01
49.82
80.63
77.50
71.29
69.11
71.06
Gemini-3-Flash
78.53
50.06
80.87
78.00
71.35
69.38
71.37
Open-source models
Table 1: Performance comparison on visual perception and reasoning benchmarks. We report accuracy (%). Δ denotes the gain over the corresponding Qwen3-VL base model.
Methods
Fine-Grained Visual Tasks
General Reasoning Tasks
Average
V* Bench
HR-Bench-4K
HR-Bench-8K
MMStar
MMBench
MathVista
Qwen3-VL-4B
Base
81.68
78.50
76.25
68.73
81.22
73.70
76.68
GRPO
84.12
80.08
77.50
70.87
82.48
75.10
78.36
OPSD
85.34
81.25
78.25
71.53
82.91
75.60
79.15
VCSD
83.77
80.38
76.88
69.20
83.08
74.55
77.98
Table 2: Comparison of training methods on fine-grained visual tasks and general reasoning tasks. We report accuracy (%) for Qwen3-VL-4B and Qwen3-VL-8B.
Setting
V*
HR-4K
HR-8K
Avg.
Answer-only
86.56
81.30
77.80
81.89
Unstructured
86.91
82.13
78.25
82.43
Oracle ROI
87.96
82.80
79.00
83.25
CVED
88.48
83.50
79.50
83.83
Table 3: CVED ablation: privileged supervision under fixed HAPI routing.
Setting
V*
HR-4K
HR-8K
Avg.
Answer-only
86.56
81.30
77.80
81.89
Unstructured
86.91
82.13
78.25
82.43
Oracle ROI
87.96
82.80
79.00
83.25
CVED
88.48
83.50
79.50
83.83
Table 3: CVED ablation: privileged supervision under fixed HAPI routing.
Setting
V*
HR-4K
HR-8K
Avg.
GRPO-only
84.12
80.08
77.50
80.57
CVED-only
86.21
81.50
78.13
81.95
Uniform hybrid
87.43
82.38
78.80
82.87
HAPI
88.48
83.50
79.50
83.83
Table 4: HAPI ablation: learning rules under fixed experience construction.
Teacher
V*Bench
ZoomBench
HR-Bench-4K
HR-Bench-8K
MME-RW-EN
MME-RW-CN
Avg.
Frozen initial
87.43
51.95
83.38
80.75
68.19
69.51
73.54
EMA (ours)
88.48
53.85
83.50
79.50
68.24
69.29
73.81
Table 5: Teacher-update ablation on Qwen3-VL-4B.
Objective
V*Bench
ZoomBench
HR-Bench-4K
HR-Bench-8K
MME-RW-EN
MME-RW-CN
Avg.
Forward KL
86.91
50.89
82.25
77.80
67.10
68.10
72.18
Reverse KL
87.96
53.96
82.95
79.75
67.80
69.05
73.58
JSD
88.48
53.85
83.50
79.50
68.24
69.29
73.81
Table 6: Divergence-objective comparison on Qwen3-VL-4B with fixed CVED and HAPI.
Figure 2: Sensitivity of VISTA on Qwen3-VL-4B to (a) rollout-group size K , (b) distillation weight λ , and (c) visual artifact budget B . Points report V*Bench, HR-4K, and HR-8K accuracy. Tested settings are equally spaced along each horizontal axis.
Appendix figures & tables7 assets
Supplementary material from the paper’s appendix.
Appendix
Channel
Content
Main setting
Visual observations
selected images from the rollout group
enabled
Interaction context
producing decisions and artifact relations
enabled
Outcome feedback
trajectory status and failure category
enabled
Appendix
Table 8: Additional teacher context used by CVED. Generated final answers and reference answers are not included.
Figure 3: Student system prompt used for active visual rollouts.
Figure 4: Collective-experience teacher instruction. The teacher evaluates the student’s unchanged response under training-only visual hindsight.
Figure 5: Teacher-context serialization with visual experience, interaction context, and outcome feedback. Braced variables are instantiated for each input.
Parameter
Value
Optimization and training duration
Optimizer
AdamW
Learning rate
1×10−6
Weight decay
0.01
Optimizer steps
65
Equivalent data passes (epochs)
≈1
Appendix
Table 9: Default VISTA training configuration for both backbone sizes. Equivalent data passes are computed as optimizer steps times input questions per batch divided by training-set size.
Experience source
V*
HR-4K
HR-8K
Avg.
Self-only hindsight
85.34
82.25
77.50
81.70
Collective visual experience
88.48
83.50
79.50
83.83
Appendix
Table 10: Experience-source ablation on Qwen3-VL-4B. Accuracy (%) is averaged over three seeds; Avg. is the unweighted mean across the three benchmarks. The collective-experience row reproduces the default CVED results in Table 4 . Bold and underlined values mark the highest and lowest scores in each column, respectively.
Setting
RL subset
Distillation subset
V*
HR-4K
HR-8K
Avg.
Uniform hybrid
All
All
87.43
82.38
78.80
82.87
RL mask only
Successful
All
86.91
82.50
78.25
82.55
Distillation mask only
All
Unsuccessful
87.43
82.25
78.92
82.87
HAPI
Successful
Unsuccessful
88.48
83.50
79.50
83.83
Appendix
Table 11: Crossed routing ablation on Qwen3-VL-4B. We report three-seed mean accuracy (%) and the unweighted benchmark average. Uniform hybrid and HAPI reproduce Table 4 . Bold and underlined values mark the highest and lowest scores in each column, respectively.