Multi-path reasoning methods such as self-consistency (SC) sample K reasoning paths and choose the most frequent answer. However, their gains quickly plateau as K increases, and existing methods do not predict when this saturation will occur. We formalize multi-path LLM reasoning as a diversity combining problem from wireless communications: each path is a noisy channel observation, and the pairwise correlation of path correctness caps the design-effect effective sample size of the vote at a finite ceiling. Generalized least squares (GLS) analysis shows that, under exchangeability, the optimal symmetric linear combiner of latent embeddings is uniform, supporting majority vote as the natural default in standard SC while leaving room for weighting or pruning under heterogeneous prompt-template branches. Across 5 models and 12 benchmarks, prompt-template diversity reduces path correlation in 55 of 57 valid cells, with the strongest effect on open-ended QA. We derive an Adaptive-K rule that uses a four-path pilot to select K∗, retaining 96--103% of MV@K=32 accuracy across Math, QA, and NLU.
Figures & tables
Fig. 1: Diversity combining for multi-path LLM reasoning. The LLM generates K correlated paths zk=hkz(x)+nk ( hk : reasoning quality; nk : errors; ρ : pairwise correlation), collapsed to correctness indicators Yk . Aggregation : under exchangeability, the GLS-optimal linear combiner is uniform (Corollary IV.5 ), supporting uniform majority vote as the natural default. Diagnostic : Keffvote=K/(1+(K−1)c) saturates at 1/c ; Adaptive-K estimates K∗ from a K=4 pilot.
K
Agree
c^
Keffvote
Ceiling
% Ceil.
pˉ (%) ↑
Tokens ↓
1
–
–
1.0
–
–
78.0 ± 2.5
323
4
0.868
0.604
1.42
1.66
86%
79.0 ± 1.3
1293
8
0.866
0.593
1.55
1.69
92%
79.2 ± 0.8
2592
16
0.862
0.579
1.65
1.73
96%
79.4 ± 0.8
5179
32
0.864
0.586
1.67
1.71
98%
79.2 ± 0.4
10327
TABLE I: Effective diversity analysis for Qwen2.5-7B on GSM8K (5 seeds × 100 instances = 500 pooled). Agree is the mean pairwise agreement (2K)−1∑j<kPr(Y(j)=Y(k)) . Correctness correlation c^ is computed via ( 3 ). Keffvote is the design-effect ( 5 ); Ceiling =1/c^ . pˉ is the mean per-path accuracy, reported as per-seed mean ± std over 5 seeds; it is not the majority-vote accuracy MV@ K of Table II . The K=1 row is a single sampled path.
Keffvote at K=
Predicted MV@32
Model
c^K=4
4
8
16
32
MV@32
BB
Binom
Qwen2.5-7B
0.60
1.42
1.55
1.65
1.67
81.9%
81.1%
100.0%
Llama-3.1-8B
0.53
1.55
1.72
1.86
1.94
79.3%
77.8%
99.9%
Mistral-7B
0.45
1.71
1.89
2.00
2.13
42.4%
41.3%
20.6%
TABLE II: Saturation and beta-binomial calibration on GSM8K (Qwen2.5-7B, Llama-3.1-8B, Mistral-7B, 5 seeds × 100 instances = 500 pooled per cell). Keffvote=K/(1+(K−1)c^) ; all models reach 96 – 98% of the ceiling 1/c^ at K=32 . MV@ K denotes the binary majority-vote accuracy at K paths (§ V-A ), which is distinct from the mean per-path accuracy pˉ of Table I . BB (beta-binomial) and Binom (independence-binomial) columns are in-sample MV@32 predictions; both estimators are formally defined in the next subsection (Eq. ( 9 )).
Figure 4
Domain
Benchmark
Answer space
n
Mean Δρ (%) ↓
Mean ΔKeffvote↑
Mean ρ^SC
Mean ρ^PT
QA
TriviaQA
open text
5
− 71.5
+ 2.1
0.68
0.19
HotpotQA
open text
5
− 62.5
+ 2.0
0.55
0.17
Science/MC
ARC-C
bounded (4)
5
− 48.7
+ 1.1
0.60
0.28
MMLU
bounded (4)
5
− 34.6
+ 0.7
0.48
0.31
Commonsense
HellaSwag
bounded (4)
5
− 36.0
+ 0.9
0.45
0.25
WinoGrande
binary (2)
5
− 36.7
+ 0.7
0.60
0.35
TABLE III: Prompt-template diversity across 12 benchmarks ( K=8 , pooled over 5 seeds, n=250 per cell except six cells with 100 – 235 instances in one arm: MATH on Qwen-7B, Qwen-32B and Llama-8B, MMLU on Qwen-7B and Qwen-32B, and MBPP on Qwen-7B). Five models evaluated: Qwen2.5-0.5B/7B/32B, Llama-3.1-8B, Mistral-7B. The Answer space column gives the effective evaluated output space (post-extraction value compared against gold), used as the relevant axis for the correlation analysis. Δρ is the relative change in pairwise correlation, (ρ^PT−ρ^SC)/ρ^SC , expressed as a percentage (so −71.5 on TriviaQA means correlation drops by 71.5% of its SC value, not by 71.5 percentage points). Cells with base accuracy <2% excluded; n = number of valid models per benchmark. Each row reports the mean across n models.
Task
Model
c^K=4
K∗
MV@ K∗
MV@32
Retained
Net cost
GSM8K (Math)
Qwen-7B
0.60
6
81.6%
81.9%
100%
31%
GSM8K (Math)
Llama-8B
0.53
8
78.1%
79.3%
98%
38%
GSM8K (Math)
Mistral-7B
0.45
10
43.6%
42.4%
103% †
44%
HotpotQA (QA)
Llama-8B
0.61
6
50.2%
52.2%
96%
31%
BoolQ (NLU)
Llama-8B
0.79
4
80.2%
80.3%
100%
25%
† Retention above 100% is consistent with finite-sample variance: the paired bootstrap
TABLE IV: Adaptive-K compute savings across three domains. c^ is estimated from K=4 SC (5 seeds ×n=100 = 500 pooled for GSM8K, 3 seeds ×n=100 = 300 for HotpotQA and BoolQ). Retained = MV accuracy at K∗ as a percentage of MV at K=32 . Net cost =(K∗+4)/32 accounts for the K=4 pilot in addition to the operating-point paths.
Eq. ( 1 ) with Assumption III.1 and the equicorrelated matrix ( 2 )
✗
Proposition III.4
Gaussian decision surrogate (Assumption III.3 )
✗
Appendix
TABLE VI: Assumption dependencies of each result. “Observable” marks results computed from output-level correctness alone.
Study
Formal saturation bound
Aggregation optimality analysis
Predicts K∗
Cross-task characterization
Answer-space analysis
SC [ 1 ]
✗
✗
✗
✗
✗
CISC [ 2 ]
✗
✗
✗
✗
✗
Entropy Voting [ 22 ]
✗
✗
✗
✗
✗
Optimal Agg. [ 23 ]
✗
∼
✗
✗
✗
SoftCoT [ 4 , 11 ]
✗
✗
✗
✗
✗
Test-time scaling [ 15 ]
✗
✗
∼
✗
✗
Appendix
TABLE VII: Analytical coverage of multi-path reasoning studies. Existing methods focus on aggregation strategies, online stopping, or empirical scaling laws; this work provides an upstream correlation-based diagnostic framework. ✓ = yes, ✗ = no, ∼ = partial. Superscripts indicate the predictive mechanism: online stops sampling per query based on observed agreement/quality; difficulty predicts K from a query-difficulty mixture model fit to scaling curves; correlation predicts K∗ from a closed-form ceiling derived from inter-path correctness correlation. Compound-inference scaling [ 6 ] fits query-difficulty heterogeneity as a mixture over scaling curves. Our c^ equals the corrected between-instance variance of per-instance accuracy divided by pˉ(1−pˉ) (§ III-B ), so it summarizes the same heterogeneity in one vote-level statistic without decomposing its sources.
Domain
n
MVSC
MVPT
ΔAcc (pp)
QA
10
6.9%
6.5%
−0.4±1.8
Sci/MC
10
60.7%
59.3%
−1.5±5.4
Commonsense
10
63.6%
63.9%
+0.3±8.7
NLU
7
50.2%
52.5%
+2.3±2.1
Code
10
13.1%
9.3%
−3.9±11.2
Math
10
46.7%
45.4%
−1.3±3.8
Appendix
TABLE VIII: Per-domain MV accuracy under SC and PT ( K=8 , 5 seeds ×n=50 pooled). n = number of valid model-benchmark cells (after pˉ<2% exclusion). ± denotes within-domain standard deviation across cells.
K
Observed MV
Beta-binomial error
Independence error
4
95.8%
0.8
2.8
8
96.8%
1.2
3.1
16
96.8%
1.3
3.2
32
96.8%
1.7
3.2
Appendix
TABLE IX: Qwen3.5-9B (thinking) on GSM8K, 3 seeds × 100 instances. Observed binary MV and held-out absolute prediction errors (pp), averaged over the two instance halves.
Model
n
c^ mean
c^ std
K∗ mean
K∗ std
Qwen-7B
25
0.571
0.122
7.9
6.77
50
0.589
0.083
6.9
1.69
100
0.596
0.055
6.6
1.02
200
0.605
0.039
6.5
0.72
Llama-8B
25
0.512
0.105
8.7
3.14
50
0.522
0.068
8.1
1.53
Appendix
TABLE X: Pilot size sensitivity on GSM8K ( K=4 , 500 bootstrap resamples of instances from pooled 5-seed data).
Model
ε
K∗
MV@ K∗
Retained
Qwen-7B
0.01
10
82.4%
101%
0.025
6
81.6%
100%
0.05
5
81.4%
99%
0.1
3
79.8%
97%
Llama-8B
0.01
13
78.8%
99%
0.025
8
78.1%
98%
Appendix
TABLE XI: Sensitivity of Adaptive-K to threshold ε on GSM8K (5 seeds ×n=100 = 500 pooled). Retained = MV@ K∗ as a percentage of MV@ K=32 .