cs.CVSep 29, 2026

VISTA: Internalizing Collective Visual Experience via On-Policy Distillation for Active Multimodal Agents

Authors: Zheng Jiang, Houde Qian, Yiming Chen, Ling Li, Chaoyang Li, Yueqi Li, Yuxuan Liu, Lifeng Sun

Organizations: Tsinghua University

Abstract

Active multimodal agents use visual tools to acquire task-relevant evidence while reasoning. Although reinforcement learning samples multiple interaction trajectories per input, outcome-based objectives primarily use the group to estimate scalar advantages, leaving complementary visual discoveries underused. We introduce VISTA, which internalizes collective visual experience through on-policy distillation by turning observations from same-input rollouts into shared supervision. Collective visual experience distillation (CVED) organizes these observations with their interaction context and aligns them with individual decisions, while heterogeneity-aware policy improvement (HAPI) reinforces successful trajectories and provides experience-guided distillation for unsuccessful attempts. An experience-conditioned teacher evaluates the student's sampled response prefixes, allowing discoveries from one trajectory to guide learning in another without replacing the student's original history or generating new target trajectories. The trained agent retains its visual tools and acts using its own interaction history. VISTA achieves the strongest average performance among the evaluated active multimodal agents of comparable size and consistently outperforms same-backbone training baselines across fine-grained perception and general reasoning tasks, demonstrating the value of collective experience for active multimodal learning.

Figures & tables

Appendix figures & tables7 assets

Supplementary material from the paper’s appendix.

Appendix

Explore similar work

CardsList
  1. VISTA: A Visual Harness for Reasoning in an Interactive World

    Oct 1, 2026Qiushi Han, Keya Hu, Linlu Qiu +2Multimodal AgentsMultimodal Reasoning

  2. On-Policy Visual Evidence Distillation

    Sep 29, 2026Shaohang Wei, Feifan Song, Guangyue Peng +9Visual EvidenceOn-Policy

  3. XSkill: Continual Learning from Experience and Skills in Multimodal Agents

    Mar 12, 2026Guanyu Jiang, Zhaochen Su, Xiaoye Qu +1Unsupervised Skill DiscoveryMultimodal Agents