Power sampling has emerged as a training-free approach to LLM reasoning, eliciting capabilities comparable to reinforcement learning by sharpening the model distribution over complete responses. Despite this success, power sampling remains underexplored in large vision-language models (LVLMs). We transfer Power-SMC to LVLM decoding by defining a sequence-power target conditioned on both the image and the prompt. This direct transfer provides a strong training-free baseline, but leaves two aspects of finite-particle multimodal inference unaddressed. At the particle level, global resampling can collapse genealogies, while particle-based power sampling does not diversify trajectories through distinct visual cues in multimodal decoding, limiting exploration under a finite particle budget. At the answer level, sequence-level sharpening makes distinct reasoning trajectories compete even when they support the same answer. We introduce ReSight-SMC, a verifier-free two-stage power sampler for LVLM inference. Its first stage uses ancestry-isolated SMC islands to preserve independent trajectory families and routes a bounded set of prefix-conditioned visual scouts to prefix-relevant image regions while discouraging redundant overlap. Each scout temporarily increases attention to the image tokens and emphasizes its routed region. Exact importance correction preserves the base LVLM sequence-power target. The second stage aggregates terminal importance mass by canonical answer, powers the answer marginal, and samples an answer together with a supporting trajectory. Across four LVLM backbones and five benchmarks, ReSight-SMC achieves stronger aggregate performance than Power-SMC over both the reasoning and perception benchmark groups. Without post-training, it remains competitive in aggregate with backbone-matched models trained using reinforcement learning.
Figures & tables
Figure 1: ReSight-SMC. Stage one maintains ancestry-isolated SMC islands for the base LVLM sequence-power target. Islands test for ESS-triggered resampling at regular token intervals. At a prespecified visual checkpoint, the remaining unfinished particles continue with the base proposal, while selected scouts use prefix-conditioned routing and attention-reactivated image tokens to form scout proposals. Exact importance correction preserves the base LVLM target. Stage two aggregates terminal mass by answer, applies a finite answer power, and samples an answer with a supporting trajectory.
System
LogicVista
MathVista
MMStar-R
MMStar-P
RealWorldQA
All-data avg.
Qwen2.5-VL-3B-Instruct
Base
31.8±2.7
46.2±1.8
42.9±0.5
49.8±4.2
53.4±4.7
45.5±1.8
Low-temp. sampling
33.5±1.2
54.4±1.0
48.3±2.1
54.9±1.8
58.2±0.9
51.1±0.4
Power-SMC
34.9±1.1
54.6±0.6
47.0±1.7
53.5±1.5
56.1±1.5
50.3±0.5
ReSight-SMC (Ours)
35.8±1.1
54.2±1.1
47.2±1.2
56.3±1.2
58.8±0.5
51.3±0.5
TRACE-RL
37.9±2.7
56.4±0.8
53.7±0.4
53.2±3.1
55.5±1.9
52.8±0.8
Table 1: Pass@1 accuracy (%) across backbones and benchmarks, reported as mean ± sample standard deviation over four seeds. The all-data average is question-weighted. Bold denotes the best non-RL result within each backbone.
Variant
LogicVista
MathVista
MMStar-R
MMStar-P
RealWorldQA
All-data avg.
ReSight-SMC
43.53±2.09
71.70±0.77
65.25±1.04
63.05±0.41
70.00±0.74
65.05±0.33
w/o islands
40.40±1.06
70.68±0.73
64.10±1.76
62.55±0.82
69.97±0.67
64.01±0.63
w/o visual scouts
41.57±1.04
71.10±0.81
64.65±0.51
62.55±0.72
70.00±0.74
64.42±0.26
w/o answer power
43.02±1.46
71.13±1.19
65.13±0.99
62.05±1.34
68.04±0.51
64.26±0.33
Table 2: Pass@1 (%) for component ablations at a matched 32-particle budget. All-data averages are question-weighted.
Task
System/readout
Pass@1
Pass@4
Coverage@32
LogicVista
Power-SMC
42.8±1.0
65.2
63.1±0.8
ReSight-SMC, γ=1
43.02±1.46
66.3
77.4±1.4
ReSight-SMC, γ=2
43.53±2.09
67.9
77.4±1.4
RealWorldQA
Power-SMC
69.2±0.4
81.2
96.5±0.2
ReSight-SMC, γ=1
68.04±0.51
81.7
96.6±0.2
ReSight-SMC, γ=2
70.00±0.74
78.3
96.6±0.2
Table 3: Representative reasoning and perception behavior with Qwen2.5-VL-7B-Instruct. The γ=1 rows replay the ReSight-SMC populations without answer-level sharpening.
(α,γ)
LogicVista
MathVista
MMStar-R
MMStar-P
RealWorldQA
All-data avg.
(2,1)
43.02±1.46
71.13±1.19
65.13±0.99
62.05±1.34
68.04±0.51
64.26±0.33
(4,1)
42.58±1.20
71.25±0.66
64.08±1.05
63.10±0.89
70.59±0.56
64.62±0.35
(2,2)
43.53±2.09
71.70±0.77
65.25±1.04
63.05±0.41
70.00±0.74
65.05±0.33
Table 4: Pass@1 (%) under alternative sharpening rules. The all-data average is question-weighted. Rows sharing α replay the same token-40 populations; changing α requires separate sampling.
Appendix figures & tables10 assets
Supplementary material from the paper’s appendix.
Appendix
Dataset
Population
Effective roots ↑
Largest-root mass ↓
LogicVista
1×32 (w/o islands)
1.625±0.048
0.900±0.006
4×8 (ReSight-SMC)
1.859±0.051
0.844±0.010
MathVista
1×32 (w/o islands)
3.790±0.030
0.671±0.004
4×8 (ReSight-SMC)
4.005±0.050
0.648±0.005
MMStar-R
1×32 (w/o islands)
2.628±0.027
0.789±0.004
4×8 (ReSight-SMC)
2.832±0.033
0.754±0.008
Appendix
Table 5: Terminal genealogy on the reasoning benchmarks with Qwen2.5-VL-7B-Instruct. The two rows for each dataset share the same 32-particle budget, visual proposal, and power parameters.
Benchmark
System
Pass@1
Pass@4
Coverage@32
MathVista
Power-SMC, 1×32
70.2±0.7
80.4
81.3±0.4
MathVista
ReSight-SMC, γ=1
71.13±1.19
81.9
86.6±0.3
MathVista
ReSight-SMC, γ=2
71.70±0.77
81.3
86.6±0.3
MMStar-R
Power-SMC, 1×32
63.7±0.8
78.4
76.3±0.9
MMStar-R
ReSight-SMC, γ=1
65.13±0.99
78.9
85.5±1.0
MMStar-R
ReSight-SMC, γ=2
65.25±1.04
78.2
85.5±1.0
Appendix
Table 6: Realized pass@1, pass@4, and Coverage@32 (%) with Qwen2.5-VL-7B-Instruct. Pass@1 and Coverage@32 are mean ± sample standard deviation; pass@4 measures success across four independently seeded complete executions.
Benchmark
Active at 32 (%)
Active at 40 (%)
Scout questions (%)
LogicVista
100.0
100.0
100.0
MathVista
100.0
99.9
100.0
MMStar-R
100.0
100.0
100.0
MMStar-P
5.9
4.7
31.2
RealWorldQA
0.0
0.0
0.0
Appendix
Table 7: Generation activity and scout activation. “Scout questions” is the fraction of question–seed runs in which at least one particle activates visual scouting.
Variant
LogicVista
MathVista
MMStar-R
MMStar-P
RealWorldQA
All-data avg.
ReSight-SMC
43.53±2.09
71.70±0.77
65.25±1.04
63.05±0.41
70.00±0.74
65.05±0.33
global attention only
42.69±1.28
70.75±0.76
64.65±0.56
62.65±0.84
70.00±0.74
64.48±0.19
island-local overlap penalty
42.80±1.36
71.58±1.35
65.00±1.28
63.20±0.59
70.00±0.74
64.88±0.69
Appendix
Table 8: Pass@1 (%) for additional visual-stage ablations with Qwen2.5-VL-7B-Instruct.
Overlap coefficient
LogicVista
MathVista
MMStar-R
MMStar-P
RealWorldQA
All-data avg.
μ=0
43.14±0.71
71.38±1.76
65.03±1.35
62.65±0.77
70.00±0.74
64.80±0.66
μ=0.5
43.02±1.79
71.80±0.62
64.45±0.29
63.25±0.53
70.00±0.74
64.83±0.37
μ=1
43.53±2.09
71.70±0.77
65.25±1.04
63.05±0.41
70.00±0.74
65.05±0.33
Appendix
Table 9: Global overlap-penalty sensitivity in realized pass@1 (%) with Qwen2.5-VL-7B-Instruct. Values are mean ± sample standard deviation over four seeds; the all-data average is question-weighted.
Area exponent
LogicVista
MathVista
MMStar-R
MMStar-P
RealWorldQA
All-data avg.
ζ=0.5
42.47±1.76
70.85±0.87
64.60±1.04
62.60±0.85
70.00±0.74
64.46±0.65
ζ=0.75
43.53±2.09
71.70±0.77
65.25±1.04
63.05±0.41
70.00±0.74
65.05±0.33
ζ=1
42.13±1.76
70.48±0.58
64.63±0.90
62.75±0.75
70.00±0.74
64.34±0.38
Appendix
Table 10: Area-exponent sensitivity in realized pass@1 (%).
Scout fraction
LogicVista
MathVista
MMStar-R
MMStar-P
RealWorldQA
All-data avg.
visual-off
41.57±1.04
71.10±0.81
64.65±0.51
62.55±0.72
70.00±0.74
64.42±0.26
ρV=0.125
43.08±0.87
71.18±0.15
64.13±0.75
63.05±1.22
70.00±0.74
64.55±0.21
ρV=0.25
43.53±2.09
71.70±0.77
65.25±1.04
63.05±0.41
70.00±0.74
65.05±0.33
ρV=0.5
43.36±0.80
71.28±0.57
64.80±0.55
62.10±1.09
70.00±0.74
64.66±0.42
Appendix
Table 11: Scout-fraction sensitivity in realized pass@1 (%).
Base α=1
Split (2,2)
Seq. α=4
A (3 weak)
9/22
729/2378
243/1380
B (2 moderate)
8/22
1024/2378
512/1380
C (1 strong)
5/22
625/2378
625/1380
Mode
A
B
C
Appendix
Table 12: A two-token autoregressive example in which base sampling, split sharpening, and direct fourth-power sequence sampling have different answer modes.
Answer position
Direct (4,1)
Two-stage (2,2)
First generated token, t∗=1
71.59%
69.12%
After Final answer: , t∗=4
71.62%
69.14%
Appendix
Table 13: Expected pass@1 in the 32-particle direct-answer example with p=(0.4,0.3,0.2,0.1) and correct answer A .
System
Population
Latency (s)
Peak VRAM (GB)
Base sampling
1
4.11
15.61
Power-SMC
1×32
7.85
17.27
ReSight-SMC
4×8
10.37
17.35
Appendix
Table 14: End-to-end inference cost with Qwen2.5-VL-7B-Instruct on one RTX 5090. Latency and peak memory are question-weighted over LogicVista, MathVista, and MMStar-R for seed 0. All systems use the Transformers backend.
Inference-time power sampling via Sequential Monte Carlo (SMC) can substantially improve large language model (LLM) reasoning without requiring post-training. However, many existing SMC approaches rely on equal-weight resampling, which can aggressively prune low-weight trajectories, discarding potentially correct reasoning paths and degrading the genealogical diversity of the search space. To address this, we introduce Chopthin-Consensus Power Sampling (CCPS). Our method applies the Chopthin resampler to LLM decoding: rather than equalizing weights and forcing unnecessary particle duplication, it enforces an upper bound on the ratio between the largest and smallest weights and carries the unequal weights forward. This targeted intervention preserves a richer set of distinct reasoning paths, keeps the weighted SMC approximation unchanged in conditional expectation, and guarantees a lower bound on the post-resampling effective sample size (ESS). To fully exploit this enriched population, we employ a semantic-majority selection mechanism that merges token-identical final trajectories, clusters semantically equivalent answers, and returns the answer supported by the largest number of distinct trajectories. Evaluating across three open-weight models and five reasoning benchmarks, we show that Chopthin increases oracle coverage in 13 of 15 settings. Combined with semantic-majority selection, CCPS matches or exceeds the final-answer accuracy of the Power-SMC baseline in 14 of 15 settings, delivering absolute gains of up to 10.6 percentage points. These findings demonstrate that diversity-preserving resampling and diversity-aware selection are complementary mechanisms for training-free LLM reasoning. Code is available at github.com/MinooAhmadii/chopthin-consensus-power-sampling.
Test-time scaling (TTS) reliably improves reasoning in large language models, but whether it transfers to small open vision-language models remains unclear. We examine this on EXAMS-V, a multilingual visual multiple-choice benchmark, comparing self-consistency, describe-then-reason with PRM-guided beam search, and two post-hoc selectors across Qwen2.5-VL-7B-Instruct and Qwen3.5-4B. What matters is the conditions under which TTS runs, not the search or verification machinery. The largest factor is parseability: an early prompt format left many chains reasoning correctly yet never committing to an answer letter, which a standard answer cue and a guided repair step largely remove. A larger decoding budget removes the rest: raising the per-chain token limit from 1k to 2k recovers 3.7 pp, whereas sampling more chains (8 to 16) adds only 0.15 pp. Once chains have room to finish, elaborate methods contribute little: PRM-guided beam search trails plain self-consistency by 0.39 pp at over eight times the cost, and neither a training-free generative critic nor a trained multimodal PRM beats majority vote across both policies. The largest gain comes instead from the policy model itself (+11.4 pp). Our best configuration reaches 84.1% on the held-out ImageCLEF 2026 test split, ranking first on the Visual MCQ leaderboard.
Spiros Baxevanakis, Peng-Jian Yang
University of Amsterdam, Science Park 904, 1098 XH Amsterdam, The Netherlands
Large Vision-Language Models (LVLMs) are rapidly evolving toward true multimodal reasoning, with visual search representing a concrete instantiation of the thinking-with-images paradigm. However, LVLM visual search faces two key challenges: incompatibility among intrinsic capabilities after post-training, and interference in long multi-step reasoning contexts. To address these, we identify two novel insights. First, self-regulation between pre- and post-training LVLMs leverages the intrinsic single-step capabilities of the pre-training model to mitigate capability deterioration and long-context interference. Second, probability-based prophetic sampling, replacing naive prompting, provides a probabilistic interface where the pre-training model acts as a prophet and the post-training model selectively accepts prophetic tokens under its output distribution, preserving coherent multi-step reasoning. Building on these insights, we introduce SeProD, a self-prophetic decoding framework that leverages intrinsic single-step capabilities to enable coherent multi-step reasoning in a training-free, plug-and-play manner. Experiments show that SeProD consistently improves multiple visual-search LVLMs across all 12 splits of 4 visual search benchmarks, as well as across general VQA benchmarks, without added computational overhead, thanks to its parallel prophetic acceptance mechanism.
Zhendong He, Qiyuan Dai, Guanbin Li +2
School of Computer Science and Engineering, Sun Yat-sen University · ShanghaiTech University