Power sampling has emerged as a training-free approach to LLM reasoning, eliciting capabilities comparable to reinforcement learning by sharpening the model distribution over complete responses. Despite this success, power sampling remains underexplored in large vision-language models (LVLMs). We transfer Power-SMC to LVLM decoding by defining a sequence-power target conditioned on both the image and the prompt. This direct transfer provides a strong training-free baseline, but leaves two aspects of finite-particle multimodal inference unaddressed. At the particle level, global resampling can collapse genealogies, while particle-based power sampling does not diversify trajectories through distinct visual cues in multimodal decoding, limiting exploration under a finite particle budget. At the answer level, sequence-level sharpening makes distinct reasoning trajectories compete even when they support the same answer. We introduce ReSight-SMC, a verifier-free two-stage power sampler for LVLM inference. Its first stage uses ancestry-isolated SMC islands to preserve independent trajectory families and routes a bounded set of prefix-conditioned visual scouts to prefix-relevant image regions while discouraging redundant overlap. Each scout temporarily increases attention to the image tokens and emphasizes its routed region. Exact importance correction preserves the base LVLM sequence-power target. The second stage aggregates terminal importance mass by canonical answer, powers the answer marginal, and samples an answer together with a supporting trajectory. Across four LVLM backbones and five benchmarks, ReSight-SMC achieves stronger aggregate performance than Power-SMC over both the reasoning and perception benchmark groups. Without post-training, it remains competitive in aggregate with backbone-matched models trained using reinforcement learning.
Figures & tables
Figure 1: ReSight-SMC. Stage one maintains ancestry-isolated SMC islands for the base LVLM sequence-power target. Islands test for ESS-triggered resampling at regular token intervals. At a prespecified visual checkpoint, the remaining unfinished particles continue with the base proposal, while selected scouts use prefix-conditioned routing and attention-reactivated image tokens to form scout proposals. Exact importance correction preserves the base LVLM target. Stage two aggregates terminal mass by answer, applies a finite answer power, and samples an answer with a supporting trajectory.
System
LogicVista
MathVista
MMStar-R
MMStar-P
RealWorldQA
All-data avg.
Qwen2.5-VL-3B-Instruct
Base
31.8±2.7
46.2±1.8
42.9±0.5
49.8±4.2
53.4±4.7
45.5±1.8
Low-temp. sampling
33.5±1.2
54.4±1.0
48.3±2.1
54.9±1.8
58.2±0.9
51.1±0.4
Power-SMC
34.9±1.1
54.6±0.6
47.0±1.7
53.5±1.5
56.1±1.5
50.3±0.5
ReSight-SMC (Ours)
35.8±1.1
54.2±1.1
47.2±1.2
56.3±1.2
58.8±0.5
51.3±0.5
TRACE-RL
37.9±2.7
56.4±0.8
53.7±0.4
53.2±3.1
55.5±1.9
52.8±0.8
Table 1: Pass@1 accuracy (%) across backbones and benchmarks, reported as mean ± sample standard deviation over four seeds. The all-data average is question-weighted. Bold denotes the best non-RL result within each backbone.
Variant
LogicVista
MathVista
MMStar-R
MMStar-P
RealWorldQA
All-data avg.
ReSight-SMC
43.53±2.09
71.70±0.77
65.25±1.04
63.05±0.41
70.00±0.74
65.05±0.33
w/o islands
40.40±1.06
70.68±0.73
64.10±1.76
62.55±0.82
69.97±0.67
64.01±0.63
w/o visual scouts
41.57±1.04
71.10±0.81
64.65±0.51
62.55±0.72
70.00±0.74
64.42±0.26
w/o answer power
43.02±1.46
71.13±1.19
65.13±0.99
62.05±1.34
68.04±0.51
64.26±0.33
Table 2: Pass@1 (%) for component ablations at a matched 32-particle budget. All-data averages are question-weighted.
Task
System/readout
Pass@1
Pass@4
Coverage@32
LogicVista
Power-SMC
42.8±1.0
65.2
63.1±0.8
ReSight-SMC, γ=1
43.02±1.46
66.3
77.4±1.4
ReSight-SMC, γ=2
43.53±2.09
67.9
77.4±1.4
RealWorldQA
Power-SMC
69.2±0.4
81.2
96.5±0.2
ReSight-SMC, γ=1
68.04±0.51
81.7
96.6±0.2
ReSight-SMC, γ=2
70.00±0.74
78.3
96.6±0.2
Table 3: Representative reasoning and perception behavior with Qwen2.5-VL-7B-Instruct. The γ=1 rows replay the ReSight-SMC populations without answer-level sharpening.
(α,γ)
LogicVista
MathVista
MMStar-R
MMStar-P
RealWorldQA
All-data avg.
(2,1)
43.02±1.46
71.13±1.19
65.13±0.99
62.05±1.34
68.04±0.51
64.26±0.33
(4,1)
42.58±1.20
71.25±0.66
64.08±1.05
63.10±0.89
70.59±0.56
64.62±0.35
(2,2)
43.53±2.09
71.70±0.77
65.25±1.04
63.05±0.41
70.00±0.74
65.05±0.33
Table 4: Pass@1 (%) under alternative sharpening rules. The all-data average is question-weighted. Rows sharing α replay the same token-40 populations; changing α requires separate sampling.
Appendix figures & tables10 assets
Supplementary material from the paper’s appendix.
Appendix
Dataset
Population
Effective roots ↑
Largest-root mass ↓
LogicVista
1×32 (w/o islands)
1.625±0.048
0.900±0.006
4×8 (ReSight-SMC)
1.859±0.051
0.844±0.010
MathVista
1×32 (w/o islands)
3.790±0.030
0.671±0.004
4×8 (ReSight-SMC)
4.005±0.050
0.648±0.005
MMStar-R
1×32 (w/o islands)
2.628±0.027
0.789±0.004
4×8 (ReSight-SMC)
2.832±0.033
0.754±0.008
Appendix
Table 5: Terminal genealogy on the reasoning benchmarks with Qwen2.5-VL-7B-Instruct. The two rows for each dataset share the same 32-particle budget, visual proposal, and power parameters.
Benchmark
System
Pass@1
Pass@4
Coverage@32
MathVista
Power-SMC, 1×32
70.2±0.7
80.4
81.3±0.4
MathVista
ReSight-SMC, γ=1
71.13±1.19
81.9
86.6±0.3
MathVista
ReSight-SMC, γ=2
71.70±0.77
81.3
86.6±0.3
MMStar-R
Power-SMC, 1×32
63.7±0.8
78.4
76.3±0.9
MMStar-R
ReSight-SMC, γ=1
65.13±0.99
78.9
85.5±1.0
MMStar-R
ReSight-SMC, γ=2
65.25±1.04
78.2
85.5±1.0
Appendix
Table 6: Realized pass@1, pass@4, and Coverage@32 (%) with Qwen2.5-VL-7B-Instruct. Pass@1 and Coverage@32 are mean ± sample standard deviation; pass@4 measures success across four independently seeded complete executions.
Benchmark
Active at 32 (%)
Active at 40 (%)
Scout questions (%)
LogicVista
100.0
100.0
100.0
MathVista
100.0
99.9
100.0
MMStar-R
100.0
100.0
100.0
MMStar-P
5.9
4.7
31.2
RealWorldQA
0.0
0.0
0.0
Appendix
Table 7: Generation activity and scout activation. “Scout questions” is the fraction of question–seed runs in which at least one particle activates visual scouting.
Variant
LogicVista
MathVista
MMStar-R
MMStar-P
RealWorldQA
All-data avg.
ReSight-SMC
43.53±2.09
71.70±0.77
65.25±1.04
63.05±0.41
70.00±0.74
65.05±0.33
global attention only
42.69±1.28
70.75±0.76
64.65±0.56
62.65±0.84
70.00±0.74
64.48±0.19
island-local overlap penalty
42.80±1.36
71.58±1.35
65.00±1.28
63.20±0.59
70.00±0.74
64.88±0.69
Appendix
Table 8: Pass@1 (%) for additional visual-stage ablations with Qwen2.5-VL-7B-Instruct.
Overlap coefficient
LogicVista
MathVista
MMStar-R
MMStar-P
RealWorldQA
All-data avg.
μ=0
43.14±0.71
71.38±1.76
65.03±1.35
62.65±0.77
70.00±0.74
64.80±0.66
μ=0.5
43.02±1.79
71.80±0.62
64.45±0.29
63.25±0.53
70.00±0.74
64.83±0.37
μ=1
43.53±2.09
71.70±0.77
65.25±1.04
63.05±0.41
70.00±0.74
65.05±0.33
Appendix
Table 9: Global overlap-penalty sensitivity in realized pass@1 (%) with Qwen2.5-VL-7B-Instruct. Values are mean ± sample standard deviation over four seeds; the all-data average is question-weighted.
Area exponent
LogicVista
MathVista
MMStar-R
MMStar-P
RealWorldQA
All-data avg.
ζ=0.5
42.47±1.76
70.85±0.87
64.60±1.04
62.60±0.85
70.00±0.74
64.46±0.65
ζ=0.75
43.53±2.09
71.70±0.77
65.25±1.04
63.05±0.41
70.00±0.74
65.05±0.33
ζ=1
42.13±1.76
70.48±0.58
64.63±0.90
62.75±0.75
70.00±0.74
64.34±0.38
Appendix
Table 10: Area-exponent sensitivity in realized pass@1 (%).
Scout fraction
LogicVista
MathVista
MMStar-R
MMStar-P
RealWorldQA
All-data avg.
visual-off
41.57±1.04
71.10±0.81
64.65±0.51
62.55±0.72
70.00±0.74
64.42±0.26
ρV=0.125
43.08±0.87
71.18±0.15
64.13±0.75
63.05±1.22
70.00±0.74
64.55±0.21
ρV=0.25
43.53±2.09
71.70±0.77
65.25±1.04
63.05±0.41
70.00±0.74
65.05±0.33
ρV=0.5
43.36±0.80
71.28±0.57
64.80±0.55
62.10±1.09
70.00±0.74
64.66±0.42
Appendix
Table 11: Scout-fraction sensitivity in realized pass@1 (%).
Base α=1
Split (2,2)
Seq. α=4
A (3 weak)
9/22
729/2378
243/1380
B (2 moderate)
8/22
1024/2378
512/1380
C (1 strong)
5/22
625/2378
625/1380
Mode
A
B
C
Appendix
Table 12: A two-token autoregressive example in which base sampling, split sharpening, and direct fourth-power sequence sampling have different answer modes.
Answer position
Direct (4,1)
Two-stage (2,2)
First generated token, t∗=1
71.59%
69.12%
After Final answer: , t∗=4
71.62%
69.14%
Appendix
Table 13: Expected pass@1 in the 32-particle direct-answer example with p=(0.4,0.3,0.2,0.1) and correct answer A .
System
Population
Latency (s)
Peak VRAM (GB)
Base sampling
1
4.11
15.61
Power-SMC
1×32
7.85
17.27
ReSight-SMC
4×8
10.37
17.35
Appendix
Table 14: End-to-end inference cost with Qwen2.5-VL-7B-Instruct on one RTX 5090. Latency and peak memory are question-weighted over LogicVista, MathVista, and MMStar-R for seed 0. All systems use the Transformers backend.