We show that multimodal models possess strong reasoning abilities and that an appropriate harness can unlock their potential to solve tasks across diverse interactive environments. We introduce VISTA, a visual harness that gives a general-purpose multimodal model long-horizon vision. VISTA allows the model to directly perceive the environment through visual observations and maintains a lossless visual memory that preserves past observations in their original form. The model can actively retrieve these observations and reorganize its visual input as it reasons. On ARC-AGI-3, VISTA improves Claude Opus 5.0's Relative Human Action Efficiency score from 40.68 to a perfect 100.00, with the model completing all 25 public games using 57.4% fewer actions than first-time human participants. VISTA's simple design also allows it to extend naturally to diverse visual environments with minimal adaptation. Across three additional benchmarks covering a diverse range of visual games and puzzles, it substantially outperforms baselines using the same underlying model with minimal harnesses. Our results highlight VISTA's potential as a general-purpose visual harness for advancing multimodal agents in complex visual environments.
Figures & tables
Figure 1: Comparison of different agent designs. (a) Language agents reason over textual observations. (b) Multimodal agents encode images into visual representations only once, but these representations may be compressed, lossy, and insufficient for subsequent reasoning, potentially limiting the models’ reasoning capabilities. (c) VISTA allows the multimodal model to iteratively call visual tools to retrieve past visual observations and reorganize its visual input as it reasons. For simplicity, we show the VLM as a vision transformer (ViT) encoder followed by an LLM. The two VLM blocks represent the same underlying model.
Figure 2: VISTA trajectories on ARC-AGI-3 (left) and GameWorld (right). The harness allows the agent to directly observe the visual states, reason in natural language, and inspect its lossless visual memory at the original level of detail when needed.
Figure 3: VISTA’s key design components and per-turn pipeline. We use ARC-AGI-3 as an example. (a) Visual observations : The model receives rendered visual states. (b) Lossless visual memory : Every returned frame is stored in its original form. (c) Model-directed visual inspection : The model selects frames and regions to examine during reasoning. (d) Per-turn pipeline : The agent observes, reasons, and acts, while each returned frame is preserved in visual memory. Revisiting past observations and updating notes are optional steps directed by the model.
System
Program
Model
Effort
RHAE
Official impl. [ 5 ]
No
GPT-5.6 Sol [ 34 ]
max
13.33
Opus 5.0 †
high
40.68
Schema [ 62 ]
Yes
GPT-5.6 Sol
xhigh → max
95.35
Opus 4.8 → Fable 5
max
98.98
ewma_sv_v1.6 [ 39 ]
Yes
GPT-5.6 Sol
xhigh
98.97
Retrodict [ 11 ]
Yes
GPT-5.6 Sol
max
99.86
Table 1: System-level comparison on 25 public games of ARC-AGI-3. VISTA is, to our knowledge, the first system to achieve perfect or near-perfect performance without program synthesis.
Figure 4: Example trajectory in ARC-AGI-3. Each frame is labeled with its turn and action; ×N denotes repeated actions. Boxes mark targets, and rings mark the agent’s clicks. The agent must infer the rules: black pieces move together with their reflections across the mirror line, and all yellow targets must be covered. Level 3 introduces a movable horizontal mirror line and selectable pieces.
Figure 5: Harness configurations on ARC-AGI-3. GPT-5.6 Sol under progressively richer harness configurations, evaluated using (a) RHAE, (b) the total number of actions taken across the 25 games, and (c) the percentage of the 25 games completed. For reference, the official implementation using textual grids (hatched gray bars) achieves a reported RHAE score of 13.33 with 10,619 actions [ 7 , 34 ] . (i) The official implementation with PNG images replacing its textual grid observations. (ii) Increases the action and time limits to match our evaluation protocol (up to 2,000 actions per game) instead of the official stopping criteria. (iii) Adds a continuous conversation, using the model’s native compaction to continue past its context limit. (iv) Adds the two notes, GUIDE.md and WORKING.md , and encourages the agent to build a compact and revisable model of the game. (v) Adds lossless visual memory and inspection. (vi) Adds exact pixel readout, completing VISTA. Configurations (i)–(iv) automatically receive up to seven frames per action, following the official protocol; configurations (v)–(vi) introduce VISTA’s key visual components: the agent automatically receives only the final frame and can actively inspect its visual memory as it reasons. See Appendix A.2 for details.
Figure 6: Comparison of observation representations on ARC-AGI-3. (a) An example game state shown as a flattened text grid and a 2D image; the displayed text grid is cropped from the full 64×64 grid. (b) RHAE, total actions across the 25 games, and average token usage per game for the two observation representations.
Figure 7: Ablations on ARC-AGI-3. Left column : RHAE. Middle column : the total number of actions taken across the 25 games. Right column : the average number of tokens per game. (a): maximum context size, which sets when context compaction occurs. (b): input image scale relative to the official 64×64 frame; scales of 2× , 4× , 8× , and 16× correspond to 128, 256, 512, and 1,024 pixels per side and roughly 20, 77, 308, and 1,229 image tokens per frame, respectively.
Figure 8: Results on ARC-AGI-3 using the open-weight GLM-5.3 Flash 320B model. We compare VISTA against the official implementation using a minimal harness, as well as a VISTA variant with inspect and read_pixels removed. VISTA’s visual harness leads to a significant performance improvement.
Figure 9: 3D renderings of ARC-AGI-3. (a) Two games rendered in 3D. (b) RHAE scores and total action counts for VISTA using 2D images and 3D renderings.
Figure 10: Model-directed use of visual memory. (a) Inspect across space and time: after the board resets, the model retrieves four frames from the preceding 22 turns in a single inspect call to reconstruct the map. (b) Inspect local visual details: the model uses inspect to zoom in on three small blocks by about 13× . The enlarged views reveal which side of each block carries a purple marker, allowing the model to identify each block’s orientation.
Figure 11: Evaluation environments and results. Top : examples from the three benchmarks. GameWorld and AI GameStore are interactive, while BabyVision is static. Bottom : VISTA with GPT-5.6 Sol compared with human players and the official implementation using a minimal harness. We report success rate and progress for GameWorld, the geometric mean of human-normalized scores (human median =100 ) for AI GameStore, and accuracy for BabyVision. The dashed line marks the best previously reported result.
Appendix figures & tables4 assets
Supplementary material from the paper’s appendix.
Appendix
Claude Opus 5.0
GPT-5.6 Sol
Game
Human actions
Score
Actions
Ratio
Score
Actions
Ratio
ar25
748
100.00
256
0.34 ×
100.00
260
0.35 ×
bp35
651
100.00
521
0.80 ×
100.00
534
0.82 ×
cd82
171
100.00
112
0.65 ×
100.00
100
0.58 ×
cn04
789
100.00
225
0.29 ×
100.00
200
0.25 ×
dc22
1,228
100.00
547
0.45 ×
100.00
560
0.46 ×
Appendix
Table 2: Per-game results on the 25 public ARC-AGI-3 games. We evaluate VISTA with Claude Opus 5.0 and GPT-5.6 Sol. For each game, human actions are the sum of the per-level human baselines, each defined as the upper-median action count among first-time human players who completed that level. The action ratio is the agent’s total actions divided by this reference. Claude Opus 5.0 achieves an RHAE score of 100.00 on every game. GPT-5.6 Sol completes every game and achieves a mean RHAE score of 99.00.
Figure 12: A full VISTA trajectory on an ARC-AGI-3 task (ID BP35), one frame per action (level 1). The player begins under upward gravity and can move left or right. Clicking a green block removes it, allowing the player to clear obstacles and navigate through otherwise blocked passages. The goal is to reach the pink object. Each frame shows the state after the labeled action; rings indicate clicks. Read left to right, top to bottom.
Figure 13: A full VISTA trajectory on an ARC-AGI-3 task (ID BP35), one frame per action (level 3). Level 3 introduces an orange block that disappears when clicked and reappears when clicked again. The agent can remove it to pass through a blocked path, or restore it to prevent the player from rising into the deadly purple spikes.
Figure 14: Four VISTA trajectories beyond ARC-AGI-3 with GPT-5.6 Sol. The top row shows GameWorld games, and the bottom row shows BabyVision tasks. Top left : in Minecraft Clone, the agent inspects across space, magnifying the hotbar and ore vein to distinguish iron from stone. Top right : in Pac-Man, it inspects across time, revisiting the current and initial frames to view the full maze. Bottom left : in Connect the Lines, the agent magnifies two regions to trace a path that is difficult to follow at the supplied image resolution, then reads exact pixel values before answering. Bottom right : in Maze, it renders the board at 2× scale and then reads the wall grid as pixel values.
Active multimodal agents use visual tools to acquire task-relevant evidence while reasoning. Although reinforcement learning samples multiple interaction trajectories per input, outcome-based objectives primarily use the group to estimate scalar advantages, leaving complementary visual discoveries underused. We introduce VISTA, which internalizes collective visual experience through on-policy distillation by turning observations from same-input rollouts into shared supervision. Collective visual experience distillation (CVED) organizes these observations with their interaction context and aligns them with individual decisions, while heterogeneity-aware policy improvement (HAPI) reinforces successful trajectories and provides experience-guided distillation for unsuccessful attempts. An experience-conditioned teacher evaluates the student's sampled response prefixes, allowing discoveries from one trajectory to guide learning in another without replacing the student's original history or generating new target trajectories. The trained agent retains its visual tools and acts using its own interaction history. VISTA achieves the strongest average performance among the evaluated active multimodal agents of comparable size and consistently outperforms same-backbone training baselines across fine-grained perception and general reasoning tasks, demonstrating the value of collective experience for active multimodal learning.
Recent progress in computer vision has produced a wide range of powerful specialized models for detection, segmentation, counting, and other visual tasks. However, these models are usually optimized for isolated task formulations, making it difficult to directly support general-purpose visual intelligence, especially when a task requires complex language understanding and dense small-object perception. In this paper, we propose VisHarness, a trainable visual agent that decouples high-level perception, reasoning, and decision-making from low-level task execution. Instead of training a model to solve a specific visual task, VisHarness learns to harness a set of carefully designed heterogeneous visual experts. This paradigm preserves the general intelligence of the agent while fully leveraging the precision advantages of specialized visual models in concrete visual tasks. With only lightweight training, VisHarness learns a generalizable visual expert-harnessing policy and can solve common fundamental vision tasks under various complex conditions through multi-turn interactions with visual expert models. To enable efficient on-policy reinforcement learning training in a live environment, we introduce dynamic visual memory archiving, which mitigates the rapidly accumulating visual-token overhead caused by multi-turn interactions with visual expert models. Experiments on four representative benchmarks covering reasoning segmentation, generalized referring segmentation, dense small-object detection, and referring counting demonstrate that VisHarness substantially outperforms existing general-purpose models and achieves competitive or superior performance compared with task-specific models.
Yaowu Fan, Tao Han, Dazhao Du +2
Sun Yat-sen University · HKUST · Harbin Institute of Technology
Agentic multimodal reasoning extends passive image understanding by allowing VLMs to actively acquire and refine visual evidence through visual tool interactions. Effective visual tool use requires three capabilities: grounding tool calls in visual context, composing tools across multiple steps, and adapting reasoning to tool-returned observations. However, existing approaches largely focus on grounding within fixed tool spaces and rigid invocation patterns, leaving composition and adaptation insufficiently addressed. We present VC-Tooler, which learns visual tool use as a compositional and adaptive capability. To this end, we first build a trajectory bank through a hierarchical synthesis pipeline covering three capability levels: single-tool grounding, multi-tool composition, and diverse tool contexts and interfaces. We then train the model in two stages: a supervised cold start that establishes these capabilities, followed by reinforcement learning that encourages accurate, efficient, and context-aware visual tool use. VC-Tooler achieves state-of-the-art performance among open-source models on both general-purpose and agentic benchmarks, including 95.8% on V* and 35.3% on VTC-Bench, and shows promising transfer under richer tool settings at inference time. Project page: https://w1zheng.github.io/VC-Tooler