Multi-path reasoning methods such as self-consistency (SC) sample K reasoning paths and choose the most frequent answer. However, their gains quickly plateau as K increases, and existing methods do not predict when this saturation will occur. We formalize multi-path LLM reasoning as a diversity combining problem from wireless communications: each path is a noisy channel observation, and the pairwise correlation of path correctness caps the design-effect effective sample size of the vote at a finite ceiling. Generalized least squares (GLS) analysis shows that, under exchangeability, the optimal symmetric linear combiner of latent embeddings is uniform, supporting majority vote as the natural default in standard SC while leaving room for weighting or pruning under heterogeneous prompt-template branches. Across 5 models and 12 benchmarks, prompt-template diversity reduces path correlation in 55 of 57 valid cells, with the strongest effect on open-ended QA. We derive an Adaptive-K rule that uses a four-path pilot to select K∗, retaining 96--103% of MV@K=32 accuracy across Math, QA, and NLU.
Figures & tables
Fig. 1: Diversity combining for multi-path LLM reasoning. The LLM generates K correlated paths zk=hkz(x)+nk ( hk : reasoning quality; nk : errors; ρ : pairwise correlation), collapsed to correctness indicators Yk . Aggregation : under exchangeability, the GLS-optimal linear combiner is uniform (Corollary IV.5 ), supporting uniform majority vote as the natural default. Diagnostic : Keffvote=K/(1+(K−1)c) saturates at 1/c ; Adaptive-K estimates K∗ from a K=4 pilot.
K
Agree
c^
Keffvote
Ceiling
% Ceil.
pˉ (%) ↑
Tokens ↓
1
–
–
1.0
–
–
78.0 ± 2.5
323
4
0.868
0.604
1.42
1.66
86%
79.0 ± 1.3
1293
8
0.866
0.593
1.55
1.69
92%
79.2 ± 0.8
2592
16
0.862
0.579
1.65
1.73
96%
79.4 ± 0.8
5179
32
0.864
0.586
1.67
1.71
98%
79.2 ± 0.4
10327
TABLE I: Effective diversity analysis for Qwen2.5-7B on GSM8K (5 seeds × 100 instances = 500 pooled). Agree is the mean pairwise agreement (2K)−1∑j<kPr(Y(j)=Y(k)) . Correctness correlation c^ is computed via ( 3 ). Keffvote is the design-effect ( 5 ); Ceiling =1/c^ . pˉ is the mean per-path accuracy, reported as per-seed mean ± std over 5 seeds; it is not the majority-vote accuracy MV@ K of Table II . The K=1 row is a single sampled path.
Keffvote at K=
Predicted MV@32
Model
c^K=4
4
8
16
32
MV@32
BB
Binom
Qwen2.5-7B
0.60
1.42
1.55
1.65
1.67
81.9%
81.1%
100.0%
Llama-3.1-8B
0.53
1.55
1.72
1.86
1.94
79.3%
77.8%
99.9%
Mistral-7B
0.45
1.71
1.89
2.00
2.13
42.4%
41.3%
20.6%
TABLE II: Saturation and beta-binomial calibration on GSM8K (Qwen2.5-7B, Llama-3.1-8B, Mistral-7B, 5 seeds × 100 instances = 500 pooled per cell). Keffvote=K/(1+(K−1)c^) ; all models reach 96 – 98% of the ceiling 1/c^ at K=32 . MV@ K denotes the binary majority-vote accuracy at K paths (§ V-A ), which is distinct from the mean per-path accuracy pˉ of Table I . BB (beta-binomial) and Binom (independence-binomial) columns are in-sample MV@32 predictions; both estimators are formally defined in the next subsection (Eq. ( 9 )).
Figure 4
Domain
Benchmark
Answer space
n
Mean Δρ (%) ↓
Mean ΔKeffvote↑
Mean ρ^SC
Mean ρ^PT
QA
TriviaQA
open text
5
− 71.5
+ 2.1
0.68
0.19
HotpotQA
open text
5
− 62.5
+ 2.0
0.55
0.17
Science/MC
ARC-C
bounded (4)
5
− 48.7
+ 1.1
0.60
0.28
MMLU
bounded (4)
5
− 34.6
+ 0.7
0.48
0.31
Commonsense
HellaSwag
bounded (4)
5
− 36.0
+ 0.9
0.45
0.25
WinoGrande
binary (2)
5
− 36.7
+ 0.7
0.60
0.35
TABLE III: Prompt-template diversity across 12 benchmarks ( K=8 , pooled over 5 seeds, n=250 per cell except six cells with 100 – 235 instances in one arm: MATH on Qwen-7B, Qwen-32B and Llama-8B, MMLU on Qwen-7B and Qwen-32B, and MBPP on Qwen-7B). Five models evaluated: Qwen2.5-0.5B/7B/32B, Llama-3.1-8B, Mistral-7B. The Answer space column gives the effective evaluated output space (post-extraction value compared against gold), used as the relevant axis for the correlation analysis. Δρ is the relative change in pairwise correlation, (ρ^PT−ρ^SC)/ρ^SC , expressed as a percentage (so −71.5 on TriviaQA means correlation drops by 71.5% of its SC value, not by 71.5 percentage points). Cells with base accuracy <2% excluded; n = number of valid models per benchmark. Each row reports the mean across n models.
Task
Model
c^K=4
K∗
MV@ K∗
MV@32
Retained
Net cost
GSM8K (Math)
Qwen-7B
0.60
6
81.6%
81.9%
100%
31%
GSM8K (Math)
Llama-8B
0.53
8
78.1%
79.3%
98%
38%
GSM8K (Math)
Mistral-7B
0.45
10
43.6%
42.4%
103% †
44%
HotpotQA (QA)
Llama-8B
0.61
6
50.2%
52.2%
96%
31%
BoolQ (NLU)
Llama-8B
0.79
4
80.2%
80.3%
100%
25%
† Retention above 100% is consistent with finite-sample variance: the paired bootstrap
TABLE IV: Adaptive-K compute savings across three domains. c^ is estimated from K=4 SC (5 seeds ×n=100 = 500 pooled for GSM8K, 3 seeds ×n=100 = 300 for HotpotQA and BoolQ). Retained = MV accuracy at K∗ as a percentage of MV at K=32 . Net cost =(K∗+4)/32 accounts for the K=4 pilot in addition to the operating-point paths.
Eq. ( 1 ) with Assumption III.1 and the equicorrelated matrix ( 2 )
✗
Proposition III.4
Gaussian decision surrogate (Assumption III.3 )
✗
Appendix
TABLE VI: Assumption dependencies of each result. “Observable” marks results computed from output-level correctness alone.
Study
Formal saturation bound
Aggregation optimality analysis
Predicts K∗
Cross-task characterization
Answer-space analysis
SC [ 1 ]
✗
✗
✗
✗
✗
CISC [ 2 ]
✗
✗
✗
✗
✗
Entropy Voting [ 22 ]
✗
✗
✗
✗
✗
Optimal Agg. [ 23 ]
✗
∼
✗
✗
✗
SoftCoT [ 4 , 11 ]
✗
✗
✗
✗
✗
Test-time scaling [ 15 ]
✗
✗
∼
✗
✗
Appendix
TABLE VII: Analytical coverage of multi-path reasoning studies. Existing methods focus on aggregation strategies, online stopping, or empirical scaling laws; this work provides an upstream correlation-based diagnostic framework. ✓ = yes, ✗ = no, ∼ = partial. Superscripts indicate the predictive mechanism: online stops sampling per query based on observed agreement/quality; difficulty predicts K from a query-difficulty mixture model fit to scaling curves; correlation predicts K∗ from a closed-form ceiling derived from inter-path correctness correlation. Compound-inference scaling [ 6 ] fits query-difficulty heterogeneity as a mixture over scaling curves. Our c^ equals the corrected between-instance variance of per-instance accuracy divided by pˉ(1−pˉ) (§ III-B ), so it summarizes the same heterogeneity in one vote-level statistic without decomposing its sources.
Domain
n
MVSC
MVPT
ΔAcc (pp)
QA
10
6.9%
6.5%
−0.4±1.8
Sci/MC
10
60.7%
59.3%
−1.5±5.4
Commonsense
10
63.6%
63.9%
+0.3±8.7
NLU
7
50.2%
52.5%
+2.3±2.1
Code
10
13.1%
9.3%
−3.9±11.2
Math
10
46.7%
45.4%
−1.3±3.8
Appendix
TABLE VIII: Per-domain MV accuracy under SC and PT ( K=8 , 5 seeds ×n=50 pooled). n = number of valid model-benchmark cells (after pˉ<2% exclusion). ± denotes within-domain standard deviation across cells.
K
Observed MV
Beta-binomial error
Independence error
4
95.8%
0.8
2.8
8
96.8%
1.2
3.1
16
96.8%
1.3
3.2
32
96.8%
1.7
3.2
Appendix
TABLE IX: Qwen3.5-9B (thinking) on GSM8K, 3 seeds × 100 instances. Observed binary MV and held-out absolute prediction errors (pp), averaged over the two instance halves.
Model
n
c^ mean
c^ std
K∗ mean
K∗ std
Qwen-7B
25
0.571
0.122
7.9
6.77
50
0.589
0.083
6.9
1.69
100
0.596
0.055
6.6
1.02
200
0.605
0.039
6.5
0.72
Llama-8B
25
0.512
0.105
8.7
3.14
50
0.522
0.068
8.1
1.53
Appendix
TABLE X: Pilot size sensitivity on GSM8K ( K=4 , 500 bootstrap resamples of instances from pooled 5-seed data).
Model
ε
K∗
MV@ K∗
Retained
Qwen-7B
0.01
10
82.4%
101%
0.025
6
81.6%
100%
0.05
5
81.4%
99%
0.1
3
79.8%
97%
Llama-8B
0.01
13
78.8%
99%
0.025
8
78.1%
98%
Appendix
TABLE XI: Sensitivity of Adaptive-K to threshold ε on GSM8K (5 seeds ×n=100 = 500 pooled). Retained = MV@ K∗ as a percentage of MV@ K=32 .
Generating multiple diverse and high-quality solutions is valuable for many applications, such as code-test generation and drug discovery. However, Large Language Models (LLMs) tend to converge on a single high-confidence solution during inference, limiting exploration of alternative valid solution paths. Existing test-time methods promote diversity through tree-like search and prune semantically similar branches using response-level semantic embeddings. However, we find that such embeddings are easily confounded by linguistic and stylistic similarities, making it difficult to distinguish genuinely distinct solution paths. To address this, we introduce Answer Probing, which probes the potential answer an LLM would reach from an intermediate reasoning path. We demonstrate that the hidden states of probed answers more effectively differentiate distinct solution paths than semantic embeddings, and the perplexity of probed answers serves as a practical proxy for reasoning correctness. Based on these findings, we propose Answer Probing-Guided Tree Search (APTS), which guides the tree search by the probed answers' hidden state similarity and perplexity. Experiments on three reasoning tasks across two LLMs show that APTS consistently enhances solution diversity, demonstrating its effectiveness and robustness.
Yi Fang, Que Shen, Chengpeng Li +6
University of Science and Technology of China · Zhongguancun Academy · Alibaba Group +2
Large language models often improve problem-solving performance by generating multi-step reasoning paths, yet how to analyze the hidden states along these paths remains unclear. Existing approaches typically assign each intermediate state the final-answer correctness label and train probes across heterogeneous questions. We argue that this protocol obscures reasoning dynamics in two ways: (1) correctness prediction can exploit question-level variation rather than path quality, and (2) states aligned by absolute step indices may correspond to different functional phases of reasoning. In this work, we propose to view reasoning paths as phase-structured trajectories within fixed questions. We instantiate this view as PAIR, short for Phase-Aligned Intra-question Reasoning. PAIR samples multiple trajectories for each question, maps variable-length paths into shared relative phases based on normalized trajectory progress, and compares successful and unsuccessful trajectories only within the same question and phase. This yields phase-specific path-quality directions that better isolate path-quality signals from question-level variation. Empirically, we find that standard across-question correctness probes lose much of their predictive power under within-question evaluation, suggesting that these probes partly rely on question-level information. PAIR improves within-question trajectory ranking and Best-of-N trajectory selection across models and benchmarks. Phase-wise steering further shows that the learned directions can change generation outcomes, providing causal evidence that they capture trajectory-relevant information.
Zhenghao He, Guangzhi Xiong, Sanchit Sinha +3
Department of Computer Science, University of Virginia
Large Language Models (LLMs) explore problems through chain-of-thought, but this exploration is buried in unstructured prose. On high-stakes tasks, users cannot tell which steps are well-supported, which alternatives were seriously considered, or how the final conclusion compares to those the model discarded. We propose a framework that ensembles the reasoning structure, not just the answers, of multiple LLMs by weighted merging of Directed Acyclic Graphs (DAGs) extracted from reasoning chains. We weight each step by how many traces independently attest to it, to return "Consensus Reasoning". Across six benchmarks spanning statutory interpretation, graduate-level science, narrative multi-hop reasoning, and first-order logic, our ensemble outperforms a matched-budget majority-vote baseline, with a maximum accuracy gain of 3.1% on MuSR-MM (narrative multi-hop reasoning). On a single model, the framework matches or exceeds self-consistency at the same trace budget while additionally exposing an inspectable consensus reasoning graph. Ensemble weights correlate with LLM-judge rankings of reasoning quality at Spearman ρ=0.30-0.51, and consensus subgraphs are preferred over alternatives leading to the majority-vote answer in 54.4-65.4% of head-to-head comparisons across five of six datasets. We observe that our framework can also be used to analyze diverse reasoning perspectives for a problem.