cs.LGSep 30, 2026

Adaptive Self-Consistency: From Black-Box Sampling to Distribution-Valued Feedback

Authors: Jingkai Huang, Yunfan Zhang, Will Ma, Weihua Zhou, Zhengyuan Zhou

Organizations: Stern School of Business New York University · School of Management Zhejiang University

Abstract

Self-consistency samples many reasoning trajectories and aggregates their final answers, treating the LLM as a black box that returns one answer per trajectory. Yet the final answer of each trajectory is sampled from a softmax vector that is available from the model's log-probabilities. We refer to this as the grey-box setting in which each trajectory reveals this answer distribution rather than a single draw from it. We formulate efficient inference in this setting as sequential mode identification with distribution-valued observations: sample trajectories one at a time and stop as soon as the LLM's modal answer is identified at a prescribed confidence level. We characterize the asymptotic stopping rate of mode identification with distribution-valued observations exactly and show that it is never worse than the black-box rate. We then propose the ASC-D algorithm, a betting stopping rule that attains this asymptotic stopping rate. On MMLU-Redux, ASC-D uses 46.446.4--95.6%95.6\% fewer trajectories than answer-only adaptive self-consistency baselines and achieves the highest fixed-budget correct-certification rate across three open-source models.

Figures & tables

Appendix figures & tables6 assets

Supplementary material from the paper’s appendix.

Appendix

Explore similar work

CardsList
  1. Optimal Self-Consistency for Efficient Reasoning with Large Language Models

    Nov 15, 2025Austin Feng, Marius Alonso, Ambroise Odonnat +2Self-ConsistencyInference-Time

  2. Self-Consistency via Marginal Sharpening

    May 27, 2026Aleksei Arzhantsev, Otmane Sakhi, Nicolas ChopinInference-TimeReasoning Skills