Multimodal agents commonly generate free-form reasoning before each action. For small models, limited model capacity can result in lengthy reasoning that provides little useful guidance for action generation while incurring substantial inference cost. To address this challenge, we introduce Selection-based Structured Reasoning (SSR), a framework that reformulates reasoning as selection instead of open-ended generation. SSR represents recurring high-level reasoning as pre-specified, reusable natural-language candidates. At each turn, the model selects from these reasoning candidates based on their likelihoods given the current context, without requiring an auxiliary task head. Using pre-specified reasoning traces enables parallel scoring, where teacher-forced prefilling computes token likelihoods concurrently within and across candidates using a shared context KV cache. We evaluate SSR on seven multimodal search benchmarks using 2B and 4B models. Across multiple reinforcement learning objectives and supervised fine-tuning, SSR delivers significant efficiency gains without sacrificing task performance. SSR achieves an average success rate competitive with leading search agents of the same scale, while reducing per-turn reasoning latency by over 90% and total per-question model inference latency by 28-54%. Project page: https://zfy0314.github.io/ssr-webpage/.
Figures & tables
Figure 1: SSR significantly reduces inference latency. Left: Parallel reasoning decoding replaces autoregressive reasoning generation. Right: Across four training objectives and two model sizes, SSR achieves comparable or higher success rates while reducing mean reasoning latency per turn by more than 90% . Mean total model inference latency per question is reduced by 28 – 54% (Table 2 ).
Figure 2: Overview of Selection-based Structured Reasoning. Freeform agents generate token-by-token reasoning before each action. SSR instead selects from a library of reusable natural-language reasoning candidates. Parallel reasoning decoding scores every token in each candidate using teacher-forced prefill, and the selected candidate guides context-specific action generation.
Model
Size
MM Search
HR-MM Search
FVQA test
Simple VQA
Live VQA
MAT Search
InfoSeek
AVG
Agentic models: zero-shot (proprietary)
GPT-4o-mini ( OpenAI, 2024a )
–
38.60
26.23
50.00
50.84
31.54
80.00
42.35
45.65
GPT-4o ( OpenAI, 2024b )
–
49.12
30.16
66.34
63.67
40.09
76.67
59.55
55.09
GPT-5 ( OpenAI, 2025b )
–
52.63
38.36
62.61
70.58
56.02
84.67
55.95
60.12
Gemini-3-Flash ( Google DeepMind, 2025 )
–
62.57
41.64
64.89
67.92
48.06
82.67
61.10
61.26
GPT-5.2 ( OpenAI, 2025a )
–
66.08
48.20
68.78
78.18
65.99
80.67
65.55
67.64
Table 1: Success rate (%) on multimodal search benchmarks. SSR achieves comparable or better performance than state-of-the-art multimodal search agents of the same size. Bold denotes the best result among 4B models; underlines mark the second best. ( ∗ reproduced results).
Training Objective
Size
Success Rate (%) ↑
Reasoning Latency per Turn (s) ↓
Model Latency per Question (s) ↓
Effective Reasoning Throughput (toks/s) ↑
GRPO-freeform
4B
58.65
0.896
5.456
72.7
GRPO-SSR
4B
61.37 (+4.6%)
0.061 (-93.2%)
2.514 (-53.9%)
3029.9 (+4070.2%)
GSPO-freeform
4B
60.45
0.939
5.195
71.9
GSPO-SSR
4B
58.60 (-3.1%)
0.063 (-93.3%)
3.706 (-28.7%)
2526.2 (+3414.6%)
SAPO-freeform
4B
61.26
1.113
6.145
69.4
SAPO-SSR
4B
60.46 (-1.3%)
0.060 (-94.6%)
2.862 (-53.4%)
3215.3 (+4535.0%)
Table 2: Success rate and inference efficiency across training objectives and model sizes. SSR achieves success rates comparable to those of freeform reasoning while reducing reasoning latency per turn by over 90% and model latency per question by 28–54%. Italic percentages show relative changes from the paired freeform model.
Method
Subset SR (%) ↑
Reasoning Latency/Turn (s)
Model Latency/Question (s)
Effective Reasoning Throughput (toks/s) ↑
Mean ↓
p95 ↓
Mean ↓
p95 ↓
MMSearch-R1-4B ( Wu et al., 2025 )
51.59
1.235
1.892
5.601
7.502
73.8
SenseNova-MARS-4B ( Chng et al., 2025 )
49.57
0.960
1.533
5.289
10.397
72.5
Chain-of-Draft ( Xu et al., 2025 )
54.06
0.416
0.798
3.657
6.616
69.9
Sketch-of-Thought ( Aytes et al., 2025 )
56.07
0.543
1.054
3.487
7.048
71.7
Efficiency-reward RL ( Arora and Zanette, 2025 )
56.28
0.756
1.273
4.638
8.675
68.9
Table 3: Success rate and efficiency measurement on the profiling subset. Our approach yields the highest success rate while being significantly faster than the baselines.
Reasoning Method
MM Search
HR-MM Search
FVQA test
Simple VQA
Live VQA
MAT Search
InfoSeek
AVG
Parallel reasoning decoding
Entire reasoning
61.40
30.16
70.22
71.08
57.19
78.67
60.90
61.37
Index only
60.82
22.30
67.56
69.99
55.56
77.33
59.30
58.98
Autoregressive decoding
With SFT
61.40
29.84
69.22
69.10
56.97
72.00
61.55
60.01
Without SFT ( ∗ degraded to freeform)
60.23
33.11
67.28
68.41
54.67
79.33
58.60
60.23
Table 4: Comparison between different reasoning methods. Selecting only the index degrades performance, while autoregressive decoding without SFT fails to limit its reasoning to the reasoning candidates in the library.
Figure 7
Appendix figures & tables3 assets
Supplementary material from the paper’s appendix.
Appendix
Subset
Samples
Accuracy reward
FVQA-train
1,894
LLM judge
DeepEyes-4K (non-MCQ)
425
LLM judge
Visual-Probe-train
912
LLM judge
DeepEyes-4K (MCQ)
628
Exact match
Total
3,859
–
Appendix
Table 5: Multimodal-search RL training data.
Name
Reasoning candidate
general-knowledge
I recognize the main subject of this image and I already know the answer from general knowledge, so I can answer the question directly.
answer-from-image
The information needed to answer is clearly visible in the image itself, so I can determine the answer by reading it directly from what I see.
examine-image-detail
The part of the image that matters for this question is small or hard to make out, so I should examine that region more closely before I decide.
identify-by-image
I cannot confidently tell what the specific entity in this image is from its appearance alone, so I should identify it by searching with the image itself before answering.
lookup-fact
I can tell what the entity in the image is, but I do not know the particular fact the question is asking about, so I should look that information up.
answer-from-evidence
The details I have gathered so far already contain what the question is asking for, so I can now give the final answer.
Appendix
Table 6: The six-candidate multimodal-search reasoning library.
Figure 5: Success rate (%) of the three reasoning types on the multimodal search benchmarks.
Agentic multimodal models have garnered significant attention for their ability to leverage external tools to tackle complex tasks. However, it is observed that such agents often meet premature interaction collapse, caused by two primary reasons: 1) the terminal reward often appending on the last token prevents the advantage from distinguishing trajectories with exploratory behavior; 2) excessively redundant context hinders the agent from absorbing useful feedback. To address these issues, we propose the Deepening Reasoning MMSearchAgent, the framework leverages the structural proximity to derive advantage signals from the whole rollout trajectories in an entire batch, such that trajectories of different lengths are further encouraged to be generated, even when containing the same correct answer. Additionally, differentiated gaussian rewards are employed to dynamically calibrate interaction tolerance, thereby ensuring information reliability and reduce redundancy. To support multi-turn interaction training, we have constructed a multi-step deep-reasoning dataset including 3602 high-quality QA pair with at least 3 reasonning steps. Extensive experiments demonstrate that our method achieves state-of-the-art performance, outperforming the MMSearch-R1 by 8.4% on FVQA-test.
Shengqin Wang, Wentao Yan, Huichi Zhou +4
East China Normal University · University College London · Huawei Noah’s Ark Lab +2
We present SimpleSearch-VL, an efficient, reliable, and practical framework for multimodal agentic search. Its core idea is to improve the agent's own search-and-verification process rather than scaling data, tools, or auxiliary model components. For efficiency, Factorized Adaptive Rollout (FAR) improves sampling efficiency by forming more informative training groups while using redundant samples to mitigate long-tail latency and expose hard samples. For reliability, SimpleSearch-VL performs evidence-verified reasoning, explicitly using chain-of-thought verification to assess the relevance of retrieved visual and textual cues to the original context. For practicality, SimpleSearch-VL keeps a lightweight tool interface and performs webpage self-summary within the agent, requiring no additional external dependencies. With only 5K supervised tool-interleaved trajectories and 2K RL data, SimpleSearch-VL improves Qwen3-VL agentic baselines by 15.8 and 16.0 average points for the 8B and 30B-A3B variants, respectively. The SimpleSearch-VL-30B-A3B model further achieves performance competitive with agentic Gemini-3-Pro.
Training multimodal agents via reinforcement learning for knowledge-intensive visual reasoning is fundamentally hindered by the extreme sparsity of outcome-based supervision and the unpredictability of live web environments. To resolve these algorithmic and environmental bottlenecks, we introduce ProMMSearchAgent, establishing a novel Sim-to-Real training paradigm for multimodal search. We decouple policy learning into a deterministic, local static sandbox. Crucially, to learn effectively within this constrained environment, we propose an introspective process-oriented reward. By probing the agent's own parametric knowledge boundaries, we generate dense behavioral metadata that explicitly rewards the correct cognitive decision, initiating a multimodal or text search only when visually or factually uncertain. Extensive experiments demonstrate that our locally-trained policy transfers zero-shot to the live Google Search API. ProMMSearchAgent achieves new SOTA performance, outperforming MMSearch-R1 by +5.1% on FVQA-test, +6.3% on InfoSeek, and +11.3% on MMSearch.
Wentao Yan, Shengqin Wang, Huichi Zhou +4
East China Normal University · Huawei Noah’s Ark Lab · University College London +2