cs.AISep 30, 2026

Learning to Route in Visual Space via Multi-Step Embedding Retrieval

Authors: Tianyu Chen, Mingyuan Zhou, Jiaxing Wu

Organizations: The University of Texas at Austin · Google DeepMind

Abstract

LLM agents rely on retrieval tools to access external knowledge, yet visual agentic search remains severely bottlenecked by standard single-step retrievers. In current pipelines, the agent must issue text queries for every intermediate step, struggling when visual clues are difficult to describe or when the retriever fails to surface necessary intermediate evidence within its top results. We hypothesize that offloading multi-step navigation across the entire embedding space directly to the retrieval tool resolves this performance bottleneck. To study this systematically, we introduce VHOP, a flexible data generation framework and benchmark with five core difficulty levels testing both visual matching and search planning. Using this framework, we develop VHOP-Router, an end-to-end training pipeline---combining supervised fine-tuning, online imitation learning, and reinforcement learning---that transforms a standard embedding model into an autoregressive multi-step retriever. Operating directly in the visual latent space, VHOP-Router retrieves linked image chains in a single tool call without requiring the agent to formulate intermediate text queries. Experiments show VHOP-Router boosts retrieval performance from under 5% to 76.3%. In agentic search, it improves task success rates by 52.7% and reduces the average token length by 61% from 1886 to 728, whereas upgrading the agent yields only a 3.7% gain. Compared to a strong baseline where the agent retrieves the top 50 results per step, VHOP-Router maintains superior performance while reducing in-context images by 23×23\times and cutting the cumulative API payload by 35×35\times. The models also generalize robustly to unseen difficulty levels and realistic test sets. Ultimately, VHOP and VHOP-Router provide an efficient and effective solution for visual agentic search that leaves native LLM capabilities entirely intact.

Figures & tables

Appendix figures & tables21 assets

Supplementary material from the paper’s appendix.

Appendix

Explore similar work

CardsList
  1. SimpleSearch-VL: A Simple Recipe for Multimodal Agentic Deep Search

    Jun 30, 2026Ming Dai, Zhihong Lu, Jinjie Gu +5Multimodal Search AgentsQwen3

  2. WeAgent-MMSearch: Native Text-Vision Interaction for Multimodal Search Agents

    Aug 28, 2026Zongkai Liu, Hui Zhang, Liqiang Niu +7Multimodal Search Agents

  3. Visual-Seeker: Towards Visual-Native Multimodal Agentic Search via Active Visual Reasoning

    Jun 13, 2026Zhengbo Zhang, Changtao Miao, Jinbo Su +10Multimodal Search AgentsCross-Modal