Multimodal agents commonly generate free-form reasoning before each action. For small models, limited model capacity can result in lengthy reasoning that provides little useful guidance for action generation while incurring substantial inference cost. To address this challenge, we introduce Selection-based Structured Reasoning (SSR), a framework that reformulates reasoning as selection instead of open-ended generation. SSR represents recurring high-level reasoning as pre-specified, reusable natural-language candidates. At each turn, the model selects from these reasoning candidates based on their likelihoods given the current context, without requiring an auxiliary task head. Using pre-specified reasoning traces enables parallel scoring, where teacher-forced prefilling computes token likelihoods concurrently within and across candidates using a shared context KV cache. We evaluate SSR on seven multimodal search benchmarks using 2B and 4B models. Across multiple reinforcement learning objectives and supervised fine-tuning, SSR delivers significant efficiency gains without sacrificing task performance. SSR achieves an average success rate competitive with leading search agents of the same scale, while reducing per-turn reasoning latency by over 90% and total per-question model inference latency by 28-54%. Project page: https://zfy0314.github.io/ssr-webpage/.
Figures & tables
Figure 1: SSR significantly reduces inference latency. Left: Parallel reasoning decoding replaces autoregressive reasoning generation. Right: Across four training objectives and two model sizes, SSR achieves comparable or higher success rates while reducing mean reasoning latency per turn by more than 90% . Mean total model inference latency per question is reduced by 28 – 54% (Table 2 ).
Figure 2: Overview of Selection-based Structured Reasoning. Freeform agents generate token-by-token reasoning before each action. SSR instead selects from a library of reusable natural-language reasoning candidates. Parallel reasoning decoding scores every token in each candidate using teacher-forced prefill, and the selected candidate guides context-specific action generation.
Model
Size
MM Search
HR-MM Search
FVQA test
Simple VQA
Live VQA
MAT Search
InfoSeek
AVG
Agentic models: zero-shot (proprietary)
GPT-4o-mini ( OpenAI, 2024a )
–
38.60
26.23
50.00
50.84
31.54
80.00
42.35
45.65
GPT-4o ( OpenAI, 2024b )
–
49.12
30.16
66.34
63.67
40.09
76.67
59.55
55.09
GPT-5 ( OpenAI, 2025b )
–
52.63
38.36
62.61
70.58
56.02
84.67
55.95
60.12
Gemini-3-Flash ( Google DeepMind, 2025 )
–
62.57
41.64
64.89
67.92
48.06
82.67
61.10
61.26
GPT-5.2 ( OpenAI, 2025a )
–
66.08
48.20
68.78
78.18
65.99
80.67
65.55
67.64
Table 1: Success rate (%) on multimodal search benchmarks. SSR achieves comparable or better performance than state-of-the-art multimodal search agents of the same size. Bold denotes the best result among 4B models; underlines mark the second best. ( ∗ reproduced results).
Training Objective
Size
Success Rate (%) ↑
Reasoning Latency per Turn (s) ↓
Model Latency per Question (s) ↓
Effective Reasoning Throughput (toks/s) ↑
GRPO-freeform
4B
58.65
0.896
5.456
72.7
GRPO-SSR
4B
61.37 (+4.6%)
0.061 (-93.2%)
2.514 (-53.9%)
3029.9 (+4070.2%)
GSPO-freeform
4B
60.45
0.939
5.195
71.9
GSPO-SSR
4B
58.60 (-3.1%)
0.063 (-93.3%)
3.706 (-28.7%)
2526.2 (+3414.6%)
SAPO-freeform
4B
61.26
1.113
6.145
69.4
SAPO-SSR
4B
60.46 (-1.3%)
0.060 (-94.6%)
2.862 (-53.4%)
3215.3 (+4535.0%)
Table 2: Success rate and inference efficiency across training objectives and model sizes. SSR achieves success rates comparable to those of freeform reasoning while reducing reasoning latency per turn by over 90% and model latency per question by 28–54%. Italic percentages show relative changes from the paired freeform model.
Method
Subset SR (%) ↑
Reasoning Latency/Turn (s)
Model Latency/Question (s)
Effective Reasoning Throughput (toks/s) ↑
Mean ↓
p95 ↓
Mean ↓
p95 ↓
MMSearch-R1-4B ( Wu et al., 2025 )
51.59
1.235
1.892
5.601
7.502
73.8
SenseNova-MARS-4B ( Chng et al., 2025 )
49.57
0.960
1.533
5.289
10.397
72.5
Chain-of-Draft ( Xu et al., 2025 )
54.06
0.416
0.798
3.657
6.616
69.9
Sketch-of-Thought ( Aytes et al., 2025 )
56.07
0.543
1.054
3.487
7.048
71.7
Efficiency-reward RL ( Arora and Zanette, 2025 )
56.28
0.756
1.273
4.638
8.675
68.9
Table 3: Success rate and efficiency measurement on the profiling subset. Our approach yields the highest success rate while being significantly faster than the baselines.
Reasoning Method
MM Search
HR-MM Search
FVQA test
Simple VQA
Live VQA
MAT Search
InfoSeek
AVG
Parallel reasoning decoding
Entire reasoning
61.40
30.16
70.22
71.08
57.19
78.67
60.90
61.37
Index only
60.82
22.30
67.56
69.99
55.56
77.33
59.30
58.98
Autoregressive decoding
With SFT
61.40
29.84
69.22
69.10
56.97
72.00
61.55
60.01
Without SFT ( ∗ degraded to freeform)
60.23
33.11
67.28
68.41
54.67
79.33
58.60
60.23
Table 4: Comparison between different reasoning methods. Selecting only the index degrades performance, while autoregressive decoding without SFT fails to limit its reasoning to the reasoning candidates in the library.
Figure 7
Appendix figures & tables3 assets
Supplementary material from the paper’s appendix.
Appendix
Subset
Samples
Accuracy reward
FVQA-train
1,894
LLM judge
DeepEyes-4K (non-MCQ)
425
LLM judge
Visual-Probe-train
912
LLM judge
DeepEyes-4K (MCQ)
628
Exact match
Total
3,859
–
Appendix
Table 5: Multimodal-search RL training data.
Name
Reasoning candidate
general-knowledge
I recognize the main subject of this image and I already know the answer from general knowledge, so I can answer the question directly.
answer-from-image
The information needed to answer is clearly visible in the image itself, so I can determine the answer by reading it directly from what I see.
examine-image-detail
The part of the image that matters for this question is small or hard to make out, so I should examine that region more closely before I decide.
identify-by-image
I cannot confidently tell what the specific entity in this image is from its appearance alone, so I should identify it by searching with the image itself before answering.
lookup-fact
I can tell what the entity in the image is, but I do not know the particular fact the question is asking about, so I should look that information up.
answer-from-evidence
The details I have gathered so far already contain what the question is asking for, so I can now give the final answer.
Appendix
Table 6: The six-candidate multimodal-search reasoning library.
Figure 5: Success rate (%) of the three reasoning types on the multimodal search benchmarks.