LLM agents rely on retrieval tools to access external knowledge, yet visual agentic search remains severely bottlenecked by standard single-step retrievers. In current pipelines, the agent must issue text queries for every intermediate step, struggling when visual clues are difficult to describe or when the retriever fails to surface necessary intermediate evidence within its top results. We hypothesize that offloading multi-step navigation across the entire embedding space directly to the retrieval tool resolves this performance bottleneck. To study this systematically, we introduce VHOP, a flexible data generation framework and benchmark with five core difficulty levels testing both visual matching and search planning. Using this framework, we develop VHOP-Router, an end-to-end training pipeline---combining supervised fine-tuning, online imitation learning, and reinforcement learning---that transforms a standard embedding model into an autoregressive multi-step retriever. Operating directly in the visual latent space, VHOP-Router retrieves linked image chains in a single tool call without requiring the agent to formulate intermediate text queries. Experiments show VHOP-Router boosts retrieval performance from under 5% to 76.3%. In agentic search, it improves task success rates by 52.7% and reduces the average token length by 61% from 1886 to 728, whereas upgrading the agent yields only a 3.7% gain. Compared to a strong baseline where the agent retrieves the top 50 results per step, VHOP-Router maintains superior performance while reducing in-context images by 23× and cutting the cumulative API payload by 35×. The models also generalize robustly to unseen difficulty levels and realistic test sets. Ultimately, VHOP and VHOP-Router provide an efficient and effective solution for visual agentic search that leaves native LLM capabilities entirely intact.
Figures & tables
Figure 1 : Task performance and upgrade gains. Agent denotes Gemini 3.5 Flash. (a) Success rate versus generated tokens plus retrieval hops. Standalone Qwen3 uses five greedy retrieval steps. (b) Gains over Agent + Qwen3 from upgrading the agent model to Gemini 3.1 Pro or replacing the retrieval tool with VHop-Router , keeping the other component fixed.
Figure 2 : Visual retrieval bottlenecks. (a) Exact symbols can be hard to describe. (b) The caption mentions “metal tools” but does not name the boxed pliers. (c) Following the green sewing machine requires the ground-truth next image. The retriever ranks it 39th, so the red piggy bank needed for the following hop is not shown.
Figure 3 : Overview of VHop and VHop-Router . Left: An L4 search example with precise and vague prompts, dead ends, and backtracking, followed by benchmark construction. Right: A complete L1 retrieval chain and small image comparisons for L2, L3, and L5, with the VHop-Router training stages below. Boxes, arrows, and detail crops are reader annotations. Examples of all difficulty levels appear in Figure 9 in Appendix B .
Split
Worlds
Queries/ world
Images/ world
Gold trajectories
SFT/ Online IL
1
978
4,492
Yes
RLVR
10
91–100
445–450
No
Table 2 : Data allocation. Query counts use precise three-hop prompts.
Figure 4 : Wider retrieval does not outperform VHop-Router . Gemini 3.5 Flash uses precise three-hop L4 test queries; VHop-Router uses precise mixed-hop training. Left: Success rate; dashed line: Flash with VHop-Router . Middle: Returned images per question. Right: Mean calls, returned images ( Imgs ), and image occurrences across agent requests ( Sent ), per question.
Figure 6 : Robustness to a foreground distractor. Left: A standard VHop image. Middle: A variant with a third, unseen distractor. Right: Retrieval success with distractor.
Figure 7 : A realistic retrieval task. Dashed insets show corpus distractors depicting the same object types but different physical instances. Zoom in for details.
Figure 8 : RLVR length-penalty dynamics. (a) Mean ∣ρ∣ . (b) L4 test success rate.
Appendix figures & tables21 assets
Supplementary material from the paper’s appendix.
Appendix
Level
Added requirement
Required behavior
L1
Object identity
Match the same physical object across views.
L2
Spatial relations
Reject an otherwise matching scene with the wrong relation.
L3
Color constraints
Check the color specified in the instruction.
L4
Dead ends
Recover after a plausible branch fails to continue.
L5
Appearance twins
Distinguish instances with similar type, color, and appearance.
L6
Shared logos
Follow links defined by a shared logo.
Appendix
Table 8 : Cumulative difficulty levels. Each level retains the requirements of the preceding levels.
Figure 9 : Visual examples of the difficulty levels. L1 shows a complete retrieval chain. L2 and L3 add spatial and color constraints; L4 introduces dead ends; and L5 requires distinguishing look-alike objects. L6 and L7 extend the matching rule to logos and printed text. The examples come from separate questions. Boxes, arrows, and enlarged details are reader annotations.
Figure 10 : State-encoder warm-up stabilizes SFT.
Figure 11 : Training dynamics of RLVR. (a) Training reward and test accuracy. (b) Rollout length and number of backtracks. (c) The fraction of rollouts that exceed budget.
Sampled action at
Basic term ℓtPG
Select (a)
−Aglogπθsel(a∣st,Ct)
Backtrack
−Aglogptbt
Stop
−Aglogptstop
Appendix
Table 9 : Basic RL policy-gradient form.
Retrieval tool
Agent
L1
L2
L3
L4
L5
L6
L7
mean
GE2 (untrained)
Flash
60.6
58.2
56.7
19.3
15.0
6.4
12.6
32.7
Pro
93.9
89.8
90.7
37.0
40.0
11.7
11.6
53.5
Qwen3 (untrained)
Flash
73.7
69.4
77.3
32.7
28.0
9.6
14.7
43.6
Pro
78.8
77.6
90.7
36.3
34.0
9.6
18.9
49.4
VHop-Router , 3 -hop
Flash
85.9
62.2
58.8
59.7
57.0
12.8
6.3
49.0
Pro
82.8
64.3
72.2
64.7
61.0
16.0
6.3
52.5
Appendix
Table 12 : Agent model and retrieval tool. Three-hop answer accuracy (%) with Gemini 3.5 Flash or Gemini 3.1 Pro, up to 16 tool calls, and the same answer-scoring convention. Both Router tools use precise training prompts. L4 uses the 300 -query test set; other levels use development world w0 . The mean averages the seven levels. Bold marks the highest value in each column.
Figure 12 : Agent and tool upgrades across all seven levels. Full version of Figure 1 b, showing gains over Flash + Qwen3 on precise three-hop queries, including both Router training configurations. Router labels indicate training hops. L4 uses the 300 -query test set; other levels use development world w0 . Gains are computed before rounding.
Retrieval tool
k
L1
L2
L3
L4
L5
L6
L7
calls
images
sent
GE2 (untrained)
1
60.6
58.2
56.7
19.3
15.0
6.4
12.6
14.5
15
120
3
82.8
81.6
78.4
46.7
40.0
10.6
24.2
11.8
35
249
5
86.9
86.7
84.5
61.0
60.0
19.1
36.8
9.6
48
300
10
92.9
93.9
88.7
78.0
72.0
31.9
42.1
6.9
69
344
50
–
–
–
81.3
–
–
–
5.4
270
1084
Qwen3 (untrained)
1
73.7
69.4
77.3
32.7
28.0
9.6
14.7
13.0
13
104
Appendix
Table 13 : Full retrieval-width comparison. Three-hop answer accuracy (%) with Flash and up to 16 tool calls. L4 uses test worlds w1 – w3 ; other levels use development world w0 . A chain call returns the policy’s active image chain. L4 context statistics are per-query means: calls counts tool calls, images counts returned images including repeats, and sent also counts repeated images in the conversation history. At k=50 , three GE2 and four Qwen3 conversations reached the request-size limit and used the last top-1 candidate as the fallback answer. Dashes indicate settings not evaluated.
Figure 13 : Full training dynamics with a decision penalty. Runs use λ∈{0,0.02,0.05} with the same 16 -decision budget. (a) Mean decisions per rollout. (b) Unpenalized terminal reward r , defined identically across runs. (c) Success rate on the preregistered 300 -question L4 test set. (d) Fraction of sampled rollouts that issue Stop before exhausting the budget. Thin curves show per-iteration values and thick curves show moving averages in (a), (b), and (d).
λ=0
λ=0.02
Δ
in-dom.
L4 test, iter 50
41.0
45.7
+4.7
L4 test, iter 100
46.0
44.7
−1.3
L4 test, iter 200
48.0
47.7
−0.3
vague
L1
32.3
37.4
+5.1
L2
34.7
49.0
+14.3
L3
27.1
31.2
+4.1
Appendix
Table 14 : Decision-penalty results. Comparison of λ=0 and λ=0.02 with gated decoding. In-domain rows use the L4 test set. Vague rows measure stopping on a verified acceptable answer, using development world w0 except for L4, which uses the test worlds. Rephrased rows use two natural-language versions of the L4 development questions; they are separate from the human-authored dataset. Stop rate and backtracks are means over the five vague-prompt levels. Accuracy is in percent; stop rate is a fraction. Bold marks the better value. Each setting uses one seed.
Training configuration
Accuracy
Precise, 3 -hop
1.0
Precise, mixed-hop
1.0
Vague, 3 -hop
8.2
Vague, mixed-hop
1.0
Appendix
Table 15 : Four-hop answer accuracy (%). Standalone policies on L4 four-hop development queries, with a 16 -decision budget. Values use the better of the two stopping constraints described above.
Method
Training configuration
Answer accuracy
Set success
VHop-Router
Precise, 3 -hop
12.0
20.0
Precise, mixed-hop
4.0
20.0
Vague, 3 -hop
4.0
4.0
Vague, mixed-hop
14.0
14.0
Flash + VHop-Router
Precise, 3 -hop
20.0
36.0
Precise, mixed-hop
20.0
34.0
Appendix
Table 16 : Full results on realistic queries. Answer accuracy and set success (%) for standalone VHop-Router and Flash with each retrieval tool. The metric definitions are given below.
Method
Strategy / training
L6 (logo)
L7 (text)
GE2
Single-shot
2.1
3.2
GE2
Greedy
4.3
2.1
GE2
History
2.1
2.1
Qwen3
Single-shot
3.2
4.2
Qwen3
Greedy
4.3
3.2
Qwen3
History
4.3
3.2
Appendix
Table 17 : Three-hop success on logo and text matching (%). These rows extend Table 4 , with the same methods and scoring rules. Both levels use precise prompts and development world w0 . The strategy or training configuration is listed in the second column.
Figure 14 : Recovery after the same incorrect first choice. Online IL and RLVR both first select 00373 . Online IL continues along that branch and stops on a wrong answer. RLVR backtracks with pbt=0.636 , retrieves the three gold images, and stops with pstop=0.999 . Solid arrows show selections and the dashed return shows backtracking. The lower rows summarize the other methods; the untrained and SFT trajectories use five fixed retrieval steps.
Figure 15 : Untrained retrieval misses the first hop. The complete five-step greedy trajectory for 2_leftof ends at 00027 . Every retrieved image lies outside the gold chain (red frames). This decoder takes five steps without learned backtracking or stopping.
Figure 16 : SFT retrieves the first hop, then leaves the chain. The state-encoder warm-up retrieves 00076 (green frame), then selects 00281 and eventually ends at 00278 . The full five-step trajectory is shown with the same decoder as the untrained model.
Figure 17 : Flash with GE2 on the same query. All 16 tool calls are shown, read down the left column and then down the right. Each row gives the recorded query, image inputs, and retrieved image. Call 2 retrieves the first gold hop, but neither the second hop nor the answer is retrieved. The recorded outcome is incorrect.
Figure 18 : Flash with Qwen3 on the same query. The complete 16 -call trace uses the same layout as Figure 17 . None of the three gold images is retrieved. The last thumbnail shows the final tool result; the figure does not supply a separate answer declaration.
Figure 19 : Re-querying the same policy recovers a valid chain. On L4 development query 19_leftof , the initial policy call exhausts its budget and returns a chain without the answer. Flash retries the same checkpoint and query image with revised text. The eighth call retrieves 00185 → 00155 → 00173 and stops correctly. Query strings are copied from the trace. The recorded system answer is 00173 ; the trace does not contain a separate FINAL declaration.
Figure 20 : Answer selection from the same returned chain. On L4 development query 33_leftof , precise mixed-hop VHop-Router retrieves the gold chain but continues searching until its 16 -action budget runs out. Its returned stack contains six images and ends at the incorrect image 00435 . Without re-querying, Flash selects the correct third image, 00247 .
Figure 21 : Two valid endpoints under one vague instruction. Family 5 from L4 test world w2 . Gray paths show the two verified chains; colored arrows show the recorded policy decisions. Solid arrows are selections and dashed arrows are backtracks. RLVR makes four selections, backtracks once, and stops at the alternative valid answer. Online IL makes eight selections and eight backtracks, exhausting its 16 -decision budget. Both are standalone policies.