cs.AIOct 1, 2026

VISTA: A Visual Harness for Reasoning in an Interactive World

Authors: Qiushi Han, Keya Hu, Linlu Qiu, Cathy Wu, Kaiming He

Organizations: Massachusetts Institute of Technology

Abstract

We show that multimodal models possess strong reasoning abilities and that an appropriate harness can unlock their potential to solve tasks across diverse interactive environments. We introduce VISTA, a visual harness that gives a general-purpose multimodal model long-horizon vision. VISTA allows the model to directly perceive the environment through visual observations and maintains a lossless visual memory that preserves past observations in their original form. The model can actively retrieve these observations and reorganize its visual input as it reasons. On ARC-AGI-3, VISTA improves Claude Opus 5.0's Relative Human Action Efficiency score from 40.68 to a perfect 100.00, with the model completing all 25 public games using 57.4% fewer actions than first-time human participants. VISTA's simple design also allows it to extend naturally to diverse visual environments with minimal adaptation. Across three additional benchmarks covering a diverse range of visual games and puzzles, it substantially outperforms baselines using the same underlying model with minimal harnesses. Our results highlight VISTA's potential as a general-purpose visual harness for advancing multimodal agents in complex visual environments.

Figures & tables

Appendix figures & tables4 assets

Supplementary material from the paper’s appendix.

Appendix

Explore similar work

CardsList
  1. Train the Agent, Not the Expert: Learning to Harness Heterogeneous Experts for Multi-Turn Visual Reasoning

    May 28, 2026Yaowu Fan, Tao Han, Dazhao Du +2

  2. VC-Tooler: Learning Compositional and Adaptive Visual Tool Use

    Aug 3, 2026Yizheng Wu, Jiashen Hua, Bing Deng +1Multimodal AgentsVisual Evidence