From Discovery to Decision: Finite-Budget Recoverability in LLM Voting
Organizations: Stony Brook University
Abstract
Voting over multiple LLM responses is a common primitive in test-time scaling and ensemble inference. Collecting more responses can expand the candidate pool and increase the chance that a correct answer is discovered. Under a fixed call budget, a discovered answer still needs to accumulate enough support within the remaining calls to become the final plurality winner, creating a discovery-to-decision gap. In this work, we characterize this gap through the realized vote state and remaining call budget. We derive a sharp recoverability threshold and show that, as sampling proceeds, the observed candidate set can only expand while the set of reachable endpoint winners can only contract, inducing a candidate-level conversion window. Under a specified iid response law, the same state yields exact finite-horizon endpoint probabilities. We further show that merging wrong-answer identities preserves single-call correctness and cannot improve plurality accuracy, and that the effect of redistributing wrong-answer probability depends on the realized vote state. Singleton reachability yields a gold-free exact locking certificate. For a known answer universe, its first trigger is the earliest prefix at which all admissible continuations yield the same fixed-budget output. Empirically, most discovered-but-unselected correct answers lose reachability only after discovery. In a controlled Word16 study, input permutation improves raw-plurality accuracy by 21.1 points with essentially unchanged single-call correctness. Exact locking saves 28-30% of calls at a 16-call budget while preserving every fixed-budget output.
Figures & tables
Appendix figures & tables42 assets
Supplementary material from the paper’s appendix.
Appendix
| Study | Purpose | Data | Scale |
|---|---|---|---|
| Forced-choice MMLU-Pro | Illustrate recovery trajectories and conversion-window closure | 200 items, five models, two sequences | 2,000 trajectories |
| GSM-Symbolic | Prospective endpoint forecasting from a realized call-8 state | 250 items in 100 template families, five models | 5,000 forecast trajectories |
| Native generation | Test recovery dynamics under direct answer generation | 1,200 MMLU-Pro and 200 Word16 items, Qwen3.5-2B/4B | 2,800 trajectories |
| Diversification | Test representation diversity and selector interaction | 256 held-out Word16 and 256 held-out MMLU-Pro items per task | 16 calls per policy |
| Physical stopping | Measure request-time consequences of exact locking | 32 MMLU-Pro items per checkpoint | batch sizes 1 and 4 |
| Judge selector | Select after plurality’s window has closed | Call-12 states of 200 MMLU-Pro items, five pools | 2,000 states, 268 vote-infeasible |
| Forced-choice MMLU-Pro a | GSM-Symbolic | Native b | Thinking mode | Public GSM8K, MATH | |
|---|---|---|---|---|---|
| Models | Qwen3.5-0.8B/2B/4B/27B, Gemma-4-E2B, Ministral-3-3B | the five 0.8B–4B checkpoints | Qwen3.5-2B/4B | Qwen3-14B/8B | Llama-3-8B/70B-Instruct |
| Prompting | chat, nonthinking, 40-word rationale | native chat, nonthinking | native chat, nonthinking | native chat, thinking | few-shot completion (5 GSM8K, 4 MATH) |
| Temperature | 0.7 | 0.7 | 0.7 (High temp. 1.0) | 0.6 | 0.6 |
| top- / top- | 0.9 / model default | 0.9 / 20 | 0.9 / 20 | 0.95 / 20 | none / none |
| Repetition penalty | 1.05 | 1.05 | 1.05 | none | none |
| Max new tokens | 96 | 1,024 | 4,096 c | 16,384 | 512 |
| Conversion-window outcome | Count | Share of discovered (%) |
|---|---|---|
| Infeasible at discovery | 74 | 5.3 |
| Closes later, before final call | 433 | 31.3 |
| Reachable through call 15, endpoint failure | 46 | 3.3 |
| Open through endpoint, gold selected | 831 | 60.0 |
| Model | Task | Converts from call 8 | Blocked / eligible at 12 | Late feasible | Late converts |
|---|---|---|---|---|---|
| 2B | MMLU-Pro | (34/201) | |||
| 4B | MMLU-Pro | (13/109) | |||
| 2B | Word16 | (37/62) | |||
| 4B | Word16 | (14/62) |
| Cohort | Trajectories | Coverage | Accuracy | I | F | L |
| Forced-choice MMLU-Pro | 2,000 | 69.20 | 41.55 | 74 | 287 | 192 |
| GSM-Symbolic | 6,250 | 89.92 | 78.82 | 68 | 359 | 267 |
| Native MMLU-Pro, 2B | 1,200 | 83.08 | 61.75 | 32 | 131 | 93 |
| Native MMLU-Pro, 4B | 1,200 | 88.42 | 76.83 | 17 | 72 | 50 |
| Native Word16, 2B | 200 | 92.50 | 70.00 | 1 | 27 | 17 |
| Native Word16, 4B | 200 | 92.00 | 50.00 | 4 | 38 | 42 |
| Task, model | Traj. | Coverage | Accuracy | Gap [95% CI] | Failures | Recoverable | Closed | Late converts | |
|---|---|---|---|---|---|---|---|---|---|
| GSM8K, 8B | 16 | 63,500 | 98.0 | 86.0 | 11.9 | 12.2 | 92.4 | 84.9 | 66/1,284 |
| 64 | 15,875 | 99.3 | 87.0 | 12.3 | 12.4 | 98.3 | 95.6 | 0/75 | |
| 256 | 3,937 | 99.8 | 87.6 | 12.2 | 12.2 | 99.2 | 98.8 | 0/5 | |
| 1024 | 889 | 100.0 | 88.3 | 11.7 | 11.7 | 100.0 | 100.0 | 0/0 | |
| GSM8K, 70B | 16 | 63,500 | 99.2 | 96.8 | 2.4 | 2.4 | 93.9 | 93.3 | 1/108 |
| 1024 | 889 | 99.9 | 96.9 | 3.0 | 3.0 | 85.2 | 100.0 | 0/4 |
| Judge | States | Judge | Direct | Difference [95% CI] |
|---|---|---|---|---|
| Own pool, pooled | 268 | 34.23 | 14.83 | |
| Qwen3.5-0.8B | 70 | 27.68 | 10.27 | |
| Qwen3.5-2B | 55 | 39.13 | 12.23 | |
| Qwen3.5-4B | 40 | 27.42 | 21.37 | |
| Gemma-4-E2B | 42 | 40.62 | 12.50 | |
| Ministral-3-3B | 61 | 35.56 | 14.44 |
| Stratum: comparator | States | Own-pool judges | Qwen3.5-27B |
|---|---|---|---|
| Covered, feasible: plurality at 16 | 248 | ||
| Currently correct: plurality at 13 | 819 | ||
| Gold absent: own direct solving | 665 | ||
| All states: plurality at 16 | 2,000 |
| Population | Observed | Forecast | Difference | |
| Wave 1 | 235 | 0.289 | 0.311 | |
| Wave 2 | 255 | 0.271 | 0.259 | |
| Pooled | 490 | 0.280 | 0.284 | |
| Forecast | 189 | 0.063 | 0.008 | |
| Forecast | 83 | 0.169 | 0.121 | |
| Forecast | 93 | 0.333 | 0.350 |
| Comparator | Population | Difference | 95% interval |
|---|---|---|---|
| Discovery-bin constant | covered, unselected | ||
| Law-only | all forecasts | ||
| Persistence | all forecasts | ||
| Law-only | covered, unselected | ||
| Persistence | covered, unselected | ||
| MMLU-Pro-fitted logistic | all forecasts |
| Data | Law | States | Observed | Forecast | Difference | (n) | AUC | Absent | |
|---|---|---|---|---|---|---|---|---|---|
| GSM-Symbolic | 16 calls (prospective) | 490 | 0.280 | 0.284 | (125) | 0.825 | 55 | ||
| 64 calls | 490 | 0.280 | 0.293 | (128) | 0.871 | 11 | |||
| GSM8K, 8B | 16 | 6,841 | 0.215 | 0.245 | (1,492) | 0.865 | 141 | ||
| 2,000 | 6,841 | 0.215 | 0.219 | (1,145) | 0.907 | 0 | |||
| GSM8K, 70B | 16 | 1,543 | 0.180 | 0.177 | (280) | 0.948 | 4 | ||
| 2,000 | 1,543 | 0.180 | 0.179 | (293) | 0.970 | 0 |
| Forecast | Squared-error gain | Endpoint Brier | Effect squared error |
|---|---|---|---|
| Full state-based | — | 0.0146 | 0.00092 |
| Uniform wrong mass | 0.0058 | 0.0212 | 0.00670 |
| Wrong-identity shuffle | 0.0030 | 0.0189 | 0.00390 |
| Forecast bin | Inputs / states | Mean forecast | Mean observed effect |
|---|---|---|---|
| 4 / 6 | |||
| 9 / 11 | |||
| 117 / 224 | |||
| 3 / 3 | |||
| 9 / 12 |
| Direction | Predicted | Observed | |
|---|---|---|---|
| 0.25 | |||
| 0.5 | |||
| 0.25 | |||
| 0.5 |
| Comparison | Effect | 95% CI | Discordant | 98.33% CI | Holm-adjusted | Decision |
|---|---|---|---|---|---|---|
| 4B Base minus 2B Base | inconclusive | |||||
| 4B Permuted minus Base | supported | |||||
| 4B Probe minus High temp. | inconclusive |
| Generation | Raw plurality | Multiset-filtered | Borda fusion |
|---|---|---|---|
| Base | 57.42 | 89.45 | 82.81 |
| Permuted | 78.52 | 95.31 | 85.55 |
| High temp. | 74.22 | 94.53 | 90.23 |
| Task | Generation | Raw | Filtered | Borda | Single call | Coverage | Invalid |
|---|---|---|---|---|---|---|---|
| Word16 | Base | 59.64 | 89.06 | 82.55 | 33.45 | 95.31 | 0.07 |
| Permuted | 74.22 | 92.19 | 85.42 | 32.28 | 98.70 | 0.10 | |
| High temp. | 76.56 | 92.45 | 85.94 | 36.49 | 97.92 | 0.10 | |
| MMLU-Pro | Base | 73.70 | 73.70 | — | 68.72 | 84.90 | 7.24 |
| Permuted | 76.30 | 76.30 | — | 69.51 | 89.06 | 6.43 | |
| High temp. | 73.96 | 73.96 | — | 68.46 | 86.46 | 7.47 |
| Arm | Valid | Truncated | Generation failure | Invalid letter | No final box |
|---|---|---|---|---|---|
| Base, | 90.87 | 8.06 | 1.04 | 0.02 | 0.01 |
| Permuted, | 90.97 | 7.97 | 1.04 | 0.01 | 0.01 |
| Base, | 90.35 | 8.58 | 1.04 | 0.00 | 0.03 |
| Permuted, | 90.37 | 8.55 | 1.04 | 0.00 | 0.03 |
| Accuracy (%) | ||||||
|---|---|---|---|---|---|---|
| Contrast | Base | Permuted | Effect (points) | 95% interval | 90% interval | Classification |
| At | 74.74 | 75.00 | practically equivalent | |||
| At | 74.48 | 73.96 | inconclusive | |||
| Interaction | inconclusive | |||||
| Quantity | Base | Permuted | Base | Permuted |
|---|---|---|---|---|
| All-call wrong collision | 0.122 | 0.109 | 0.113 | 0.106 |
| Collision among errors | 0.693 | 0.634 | 0.676 | 0.599 |
| Distinct wrong answers per input | 0.79 | 0.95 | 0.84 | 1.00 |
| Coverage at 16 calls (%) | 86.9 | 88.0 | 85.9 | 87.2 |
| Single-call correctness (%) | 68.7 | 67.8 | 67.6 | 66.9 |
| Metric | MMLU-Pro ( ) | GSM ( ) |
|---|---|---|
| Locked by call 9 | ||
| Locked by call 12 | ||
| Not locked before call 16 | ||
| First lock, quartiles | ||
| Mean calls consumed | ||
| Calls saved |
| Cohort | Trajectory law | Checkpoints | Traj. | Obs. | Pred. | Difference [95% CI] | KS (bound) | MAD | |
|---|---|---|---|---|---|---|---|---|---|
| GSM-Symbolic | 16 | – | five a | 4,000 | 32.64 | 32.48 | 0.012 (0.022) | 0.067 | |
| GSM-Symbolic | 32 | three | 750 | 39.06 | 39.01 | 0.010 (0.028) | 0.040 | ||
| GSM-Symbolic | 32 | three | 750 | 39.06 | 39.04 | 0.011 (0.029) | 0.039 | ||
| GSM-Symbolic | 48 | three | 750 | 40.02 | 39.83 | 0.013 (0.031) | 0.033 | ||
| GSM-Symbolic | 64 | three | 750 | 40.32 | 40.20 | 0.018 (0.044) | 0.034 | ||
| MMLU-Pro | 16 | seq. 0 seq. 1 | five | 1,000 | 28.15 | 28.81 | 0.039 (0.057) | 0.092 |
| Model | Traj. | Obs. | Pred. | Difference [95% CI] | Asymptote | MAD [95% CI] | KS (bound) | Var. obs./CLT | |
|---|---|---|---|---|---|---|---|---|---|
| Llama-3-8B | 16 | 63,500 | 34.90 | 34.85 | 39.53 | 0.0565 | 0.004 (0.008) | 1.5 / 2.4 | |
| 32 | 31,750 | 37.34 | 37.29 | 0.0297 | 0.007 (0.012) | 3.4 / 4.7 | |||
| 64 | 15,875 | 38.55 | 38.48 | 0.0155 | 0.006 (0.013) | 7.5 / 9.5 | |||
| 128 | 7,874 | 39.10 | 39.03 | 0.0087 | 0.006 (0.014) | 16.5 / 18.9 | |||
| 256 | 3,937 | 39.38 | 39.29 | 0.0056 | 0.012 (0.024) | 36.1 / 37.8 | |||
| 512 | 1,905 | 39.48 | 39.42 | 0.0042 | 0.010 (0.024) | 68.0 / 75.7 |
| Model | Traj. | Obs. | Pred. | Difference [95% CI] | Asymptote | MAD [95% CI] | KS (bound) | Var. obs./CLT | |
|---|---|---|---|---|---|---|---|---|---|
| Llama-3-8B | 16 | 64,000 | 13.28 | 13.19 | 15.75 | 0.0419 | 0.006 (0.010) | 1.4 / 3.3 | |
| 64 | 16,000 | 15.17 | 15.05 | 0.0178 | 0.009 (0.018) | 9.3 / 13.1 | |||
| 256 | 3,968 | 15.70 | 15.57 | 0.0100 | 0.012 (0.029) | 46.5 / 52.3 | |||
| 1024 | 896 | 15.82 | 15.72 | 0.0088 | 0.017 (0.043) | 203.1 / 209.3 | |||
| Llama-3-70B | 16 | 64,000 | 22.97 | 22.85 | 26.25 | 0.0462 | 0.004 (0.008) | 1.5 / 3.1 | |
| 64 | 16,000 | 25.60 | 25.51 | 0.0154 | 0.005 (0.013) | 8.7 / 12.3 |
| Model | Tokens omitted | Calls omitted | Token call | Obs. pred. | KS (bound) | |
|---|---|---|---|---|---|---|
| Qwen3-14B | 32 | 41.9 | 45.1 | 0.009 (0.021) | ||
| 24 | 41.3 | 44.1 | 0.011 (0.024) | |||
| 16 | 38.9 | 41.9 | 0.008 (0.026) | |||
| Qwen3-8B | 32 | 40.6 | 44.1 | 0.011 (0.024) | ||
| 24 | 40.0 | 43.3 | 0.011 (0.029) | |||
| 16 | 38.4 | 41.3 | 0.009 (0.023) |
| Model | Rule | Calls omitted | Tokens omitted | Changed endpoints | Guarantee |
|---|---|---|---|---|---|
| Qwen3-14B | Certificate | 45.1 | 41.9 | 0 | yes |
| ASC 0.95 | 84.5 | 78.4 | 0 | no | |
| ESC, window 8 | 70.0 | 62.0 | 0 | no | |
| Qwen3-8B | Certificate | 44.1 | 40.6 | 0 | yes |
| ASC 0.95 | 83.0 | 77.1 | 0 | no | |
| ESC, window 8 | 67.0 | 58.0 | 0 | no |
| Method | Certificate | ASC | ESC | Bayesian |
|---|---|---|---|---|
| Changed outputs | 0 | 1 | 6 | 22 |
| Accuracy difference | (guaranteed) | |||
| 95% interval |
| Model | Batch | Certificate | ASC .95 | ESC4 | Bayesian |
|---|---|---|---|---|---|
| 2B | 1 | ||||
| 2B | 4 | ||||
| 4B | 1 | ||||
| 4B | 4 |
| Word16 | MMLU-Pro | Shortest path | |||||||
|---|---|---|---|---|---|---|---|---|---|
| Method | Acc. | Calls | Time | Acc. | Calls | Time | Acc. | Calls | Time |
| Fixed-16 | 95.83 | 16.00 | 100.43 | 85.42 | 16.00 | 113.09 | 93.75 | 16.00 | 342.62 |
| Certificate | 95.83 | 14.50 | 90.82 | 85.42 | 11.75 | 88.67 | 93.75 | 13.50 | 289.50 |
| ASC 0.95 | 95.83 | 13.50 | 86.19 | 85.42 | 6.33 | 58.00 | 93.75 | 10.92 | 235.03 |
| ASC 0.99 | 95.83 | 15.17 | 94.94 | 85.42 | 10.33 | 82.78 | 93.75 | 13.00 | 279.66 |
| ESC window 4 | 95.83 | 15.42 | 97.23 | 85.42 | 6.92 | 65.20 | 93.75 | 13.58 | 292.62 |
| Rule | Parameter | Mean stop | Saved % | Changed % | Accuracy change (points) |
|---|---|---|---|---|---|
| Exact certificate | 11.46 | 28.35 | 0.00 | 0.00 | |
| Majority lock | 12.44 | 22.25 | 0.00 | 0.00 | |
| Plug-in curtailment, | 7.07 | 55.80 | 2.10 | ||
| 6.12 | 61.76 | 3.50 | |||
| 4.40 | 72.51 | 7.85 | |||
| 2.65 | 83.42 | 17.60 |
| Checkpoint | Revision |
|---|---|
| Qwen/Qwen3.5-0.8B | 2fc06364 715b967f 1860aea9 cf387788 75588b17 |
| Qwen/Qwen3.5-2B | 15852e8c 16360a2f ea060d61 5a32b452 70f8a8fc |
| Qwen/Qwen3.5-4B | 851bf6e8 06efd8d0 a36b00dd f55e13cc b7b8cd0a |
| google/gemma-4-E2B-it | 3e22461f 65e89153 144f8adb 70e3b8c2 cc9845a7 |
| mistralai/Ministral-3-3B-Instruct-2512-BF16 | b6d637be f2393152 b3da2b2f de72eecd ee30557e |
| Qwen/Qwen3.5-27B (judge and forced-choice replay, Sections E.5 and E.4 ) | fc05daec 18b0a78c 049392ed 2e771dde 82bdf654 |