cs.CVSep 28, 2026

ReSight-SMC: Two-Stage Power Sampling via Island SMC with Visual Scouts

Authors: Yaowen Zhang, Xiangyu Qiu, Junyi Hu, Zhi Lu, Wenwen Tian, Aoqin Wang, Junhai Luo, Zhenming Peng

Organizations: Independent Researcher · University of Electronic Science and Technology of China · Tsinghua University

Abstract

Power sampling has emerged as a training-free approach to LLM reasoning, eliciting capabilities comparable to reinforcement learning by sharpening the model distribution over complete responses. Despite this success, power sampling remains underexplored in large vision-language models (LVLMs). We transfer Power-SMC to LVLM decoding by defining a sequence-power target conditioned on both the image and the prompt. This direct transfer provides a strong training-free baseline, but leaves two aspects of finite-particle multimodal inference unaddressed. At the particle level, global resampling can collapse genealogies, while particle-based power sampling does not diversify trajectories through distinct visual cues in multimodal decoding, limiting exploration under a finite particle budget. At the answer level, sequence-level sharpening makes distinct reasoning trajectories compete even when they support the same answer. We introduce ReSight-SMC, a verifier-free two-stage power sampler for LVLM inference. Its first stage uses ancestry-isolated SMC islands to preserve independent trajectory families and routes a bounded set of prefix-conditioned visual scouts to prefix-relevant image regions while discouraging redundant overlap. Each scout temporarily increases attention to the image tokens and emphasizes its routed region. Exact importance correction preserves the base LVLM sequence-power target. The second stage aggregates terminal importance mass by canonical answer, powers the answer marginal, and samples an answer together with a supporting trajectory. Across four LVLM backbones and five benchmarks, ReSight-SMC achieves stronger aggregate performance than Power-SMC over both the reasoning and perception benchmark groups. Without post-training, it remains competitive in aggregate with backbone-matched models trained using reinforcement learning.

Figures & tables

Appendix figures & tables10 assets

Supplementary material from the paper’s appendix.

Appendix

Explore similar work

CardsList
  1. Chopthin-Consensus Power Sampling: A Diversity-Preserving Approach to LLM Decoding

    Sep 14, 2026Minoo Ahmadi, Seyedarmin Azizi, Erfan Baghaei Potraghloo +2LLM Reasoning StrategiesLarge Language Model Decoding

  2. Test-Time Scaling for Small VLMs on Multilingual Visual MCQ

    Jul 10, 2026Spiros Baxevanakis, Peng-Jian YangTest-Time ScalingMultilingual Benchmark

  3. Self-Prophetic Decoding to Unlock Visual Search in LVLMs

    May 27, 2026Zhendong He, Qiyuan Dai, Guanbin Li +2Recent Vision-Language ModelsMultimodal Reasoning