Beyond Solo and Consistency: Vindicating Multi-Agent Debate via Conditional Progressive Pruning
Organizations: Rutgers University, New Brunswick · University of Cambridge · Tsinghua University · Harvard University · Case Western Reserve University · University of California San Diego · Carnegie Mellon University
Abstract
Large Language Model (LLM) based Multi-Agent Debate (MAD) is one of the most effective test time scaling techniques. Through multi-round communication, agents complement each other in knowledge and reasoning and solve tasks that no single member can solve. However, existing MAD frameworks fail to beat strong Single Agent and Consistency-based baselines under the same strict cost limit, which shakes the foundation of the MAD field. We propose Conditional Progressive Pruning (CPP), a lightweight pruning framework that fully exploits multi-round MAD. CPP outperforms all existing MAD frameworks on multiple dominated benchmarks. It is also the first to fully outperform consistency methods. Our code, detailed agent interaction records will be released soon.
Figures & tables
| MATH-Hard | GPQA-Diamond | MMLU-Pro | ||||
| Method | Acc (%) | Cost | Acc (%) | Cost | Acc (%) | Cost |
| Grok | 48.50 1.38 | 11.17% | 57.33 4.27 | 13.50% | 57.00 2.90 | 12.05% |
| GPT | 43.83 4.79 | 7.56% | 38.67 3.78 | 6.78% | 42.17 3.43 | 6.49% |
| Mistral | 46.17 3.97 | 8.15% | 45.33 3.44 | 4.46% | 57.33 1.51 | 5.22% |
| Vote | 55.00 3.63 | 26.88% | 58.67 3.01 | 24.74% | 60.67 3.61 | 23.75% |
| IoE-Grok | 56.83 3.06 | 28.24% | 65.83 2.32 | 32.79% | 58.83 3.60 | 32.84% |
| MATH-Hard | GPQA-Diamond | MMLU-Pro | ||||
| Method | Acc (%) | Cost | Acc (%) | Cost | Acc (%) | Cost |
| Consistency-Grok | 58.50 1.97 | 101.06% | 63.33 1.63 | 91.43% | 61.00 1.55 | 77.05% |
| Consistency-GPT | 57.67 1.03 | 84.88% | 47.83 1.60 | 68.47% | 48.50 1.05 | 71.40% |
| Consistency-Mistral | 57.50 1.97 | 99.21% | 45.33 3.44 | 4.46% | 58.83 1.72 | 16.29% |
| Consistency-Fuse | 64.00 2.45 | 107.89% | 62.67 2.16 | 56.88% | 66.17 2.32 | 78.94% |
| Consistency-Best | 62.83 2.14 | 103.48% | 63.17 2.48 | 74.85% | 64.17 3.06 | 86.30% |
| Vanilla | Adversarial | |||||||
| Method | Ref | Acc (%) | Cost | Edge | Ref | Acc (%) | Cost | Edge |
| Grok | – | 41.17 3.60 | 13.24% | – | – | 9.83 2.48 | 9.78% | – |
| GPT | – | 43.83 4.79 | 3.95% | – | – | 4.67 2.34 | 1.01% | – |
| Mistral | – | 46.17 3.97 | 4.26% | – | – | 12.67 2.07 | 1.07% | – |
| Vote | – | 51.17 3.31 | 21.45% | – | – | 10.17 1.72 | 11.86% | – |
| IoE-Vote | – | 52.83 3.97 | 51.66% | – | – | 47.17 4.26 | 52.34% | – |
| Pipeline | Reward | ||||||||
| Ablation | Ref | Acc (%) | Cost | Edge | Ablation | Ref | Acc (%) | Cost | Edge |
| 10-Sample | 0.91, 0.86, 0.86 | 66.00 2.37 | 96.73% | 15.77 | R_ + a | 0.91, 0.87, 0.86 | 65.17 3.37 | 93.79% | 15.83 |
| 30-Sample | 0.95, 0.92, 0.90 | 64.67 3.08 | 97.34% | 16.61 | R_0.5a | 0.73, 0.7, 0.69 | 62.33 2.25 | 86.91% | 12.70 |
| 90-Sample | 0.92, 0.88, 0.87 | 64.00 1.10 | 96.27% | 16.08 | R_0 | 0.53, 0.5, 0.5 | 59.50 2.66 | 81.20% | 9.17 |
| 150-Sample | 0.87, 0.85, 0.84 | 62.33 3.50 | 94.24% | 15.34 | R_ - a | 0.39, 0.39, 0.41 | 60.00 2.10 | 77.14% | 7.18 |
| w/o Training | 0.95, 0.92, 0.91 | 63.33 2.94 | 96.66% | 16.73 | w/o Round-Var. | 0.83, 0.8, 0.79 | 64.67 1.86 | 92.94% | 14.50 |
| Method | q5/ q5/ q5 | q75/ q75/ q75 | q80/ q90/ q90 | q85/ q95/ q95 |
| Full-Debate | 52.83 3.31 68.52% | 58.83 2.04 85.87% | 60.00 2.00 88.51% | 60.67 2.94 90.30% |
| ARG-Designer | 52.67 2.73 68.24% | 61.83 4.12 85.26% | 59.00 2.37 87.58% | 60.33 3.14 89.92% |
| Adv_CPP | 53.83 2.04 66.26% | 57.00 3.52 80.77% | 62.17 3.54 83.27% | 63.00 3.16 84.48% |
| IoE-Vote ∗ | 52.83 3.97 78.82% | |||
| Debate-IoE | 61.00 2.53 87.39% | |||
| Consistency-Fuse-12 | 62.00 1.67 68.46% | |||
Appendix figures & tables24 assets
Supplementary material from the paper’s appendix.
Appendix
| Action | Dependency |
| MATH-Hard | GPQA-Diamond | MMLU-Pro | ||||
| Method | Ref | Edge | Ref | Edge | Ref | Edge |
| Full-Debate | 1, 1, 1 | 18.00 | 1, 1, 1 | 18.00 | 1, 1, 1 | 18.00 |
| Random | 0.5, 0.51, 0.5 | 9.02 | 0.5, 0.5, 0.5 | 9.02 | 0.5, 0.5, 0.5 | 9.02 |
| LLM-Self | 1.67, 1.58, 1.62 | 29.24 | 1.76, 1.47, 1.51 | 28.48 | 1.77, 1.58, 1.61 | 29.80 |
| Intervention | 0.4, 0.77, 0.83 | 12.00 | 0.53, 0.75, 0.69 | 12.00 | 0.56, 0.72, 0.72 | 12.00 |
| -MAD | 0.57, 0.58, 0.57 | 10.31 | 0.52, 0.56, 0.55 | 9.77 | 0.45, 0.47, 0.46 | 8.24 |
| Threshold | Grok | GPT | Mistral |
| q5 | 65 | 132 | 139 |
| q75 | 157 | 318 | 432 |
| q80 | 166 | 353 | 478 |
| q85 | 177 | 393 | 527 |
| q90 | 196 | 458 | 623 |
| q95 | 233 | 621 | 872 |
| 2 | 3 | 4 | 5 | 6 | 7 | 8 | 9 | Cost | |
| Grok | 49.33 2.34 | 51.17 2.40 | 53.33 3.83 | 55.00 3.10 | 56.00 0.63 | 56.83 1.94 | 57.83 1.47 | 58.50 1.97 | 101.06% |
| 2 | 3 | 4 | 5 | 6 | 7 | 8 | 9 | 10 | 11 | 12 | 13 | Cost | |
| GPT | 47.67 2.25 | 49.33 2.42 | 51.67 1.51 | 53.33 2.07 | 53.50 2.43 | 54.00 1.79 | 54.67 1.03 | 55.83 0.98 | 57.33 2.34 | 57.67 1.03 | 57.00 1.55 | 57.17 2.04 | 100.32% |
| 2 | 3 | 4 | 5 | 6 | 7 | 8 | 9 | 10 | 11 | 12 | 13 | Cost | |
| Mistral | 49.00 4.24 | 52.67 2.16 | 54.50 3.78 | 56.00 2.83 | 55.83 2.32 | 56.67 2.25 | 56.33 2.42 | 56.67 1.97 | 57.17 1.83 | 56.83 1.72 | 57.50 1.97 | 57.17 1.94 | 107.48% |
| 2 | 3 | 4 | 5 | 6 | 7 | 8 | 9 | 10 | 11 | 12 | Cost | |
| Fuse-Extension | 49.83 1.83 | 52.67 1.21 | 53.83 2.64 | 55.50 2.43 | 58.67 1.03 | 59.83 1.17 | 60.50 2.59 | 62.33 1.63 | 62.83 2.14 | 63.50 3.27 | 64.00 2.45 | 107.89% |
| 2 | 3 | 4 | 5 | 6 | 7 | 8 | 9 | 10 | 11 | Cost | |
| Best-Extension | 50.83 2.32 | 53.17 1.60 | 55.00 2.10 | 57.33 2.66 | 58.17 2.71 | 59.17 2.64 | 61.33 1.51 | 62.00 2.10 | 62.50 1.52 | 62.83 2.14 | 103.48% |
| 2 | 3 | 4 | 5 | 6 | 7 | 8 | 9 | 10 | 11 | 12 | |
| Mistral | 48.50 4.85 | 51.83 3.37 | 54.33 3.14 | 56.33 2.80 | 55.50 1.97 | 56.50 2.07 | 56.67 2.42 | 57.00 2.00 | 57.00 2.45 | 56.83 1.17 | 58.00 1.79 |
| 2 | 3 | 4 | 5 | 6 | 7 | 8 | 9 | 10 | 11 | 12 | 13 | 14 | Cost | |
| Fuse-Extension | 48.00 3.35 | 51.17 2.14 | 53.33 2.73 | 53.33 2.58 | 56.50 2.26 | 58.33 2.34 | 59.17 0.98 | 60.00 1.10 | 61.83 1.47 | 62.17 1.17 | 62.00 1.67 | 62.00 1.67 | 62.50 0.84 | 94.94% |
| 2 | 3 | 4 | 5 | 6 | 7 | 8 | 9 | 10 | 11 | 12 | 13 | 14 | 15 | 16 | Cost | |
| Best-Extension | 48.00 3.90 | 51.00 2.37 | 53.33 3.98 | 53.17 2.99 | 55.83 2.14 | 56.50 2.66 | 58.17 2.79 | 57.33 3.44 | 58.50 1.97 | 60.00 1.10 | 60.50 1.64 | 60.33 1.37 | 60.83 1.60 | 61.50 1.52 | 62.17 1.72 | 94.81% |
| 2 | 3 | 4 | 5 | 6 | 7 | 8 | 9 | 10 | 11 | 12 | |
| Mistral | 11.67 4.03 | 10.50 1.52 | 13.50 2.74 | 15.17 1.60 | 15.67 3.14 | 15.50 1.87 | 17.00 2.28 | 16.67 2.73 | 17.83 1.83 | 18.00 2.76 | 19.17 1.17 |
| 2 | 3 | 4 | 5 | 6 | 7 | 8 | 9 | 10 | 11 | 12 | 13 | 14 | Cost | |
| Fuse-Extension | 12.67 2.42 | 12.50 2.59 | 12.17 2.79 | 13.83 2.64 | 14.00 2.68 | 14.83 2.64 | 14.33 2.34 | 14.17 2.48 | 14.00 1.41 | 14.67 2.50 | 16.17 1.72 | 16.33 2.80 | 18.00 1.26 | 54.67% |
| 2 | 3 | 4 | 5 | 6 | 7 | 8 | 9 | 10 | 11 | 12 | 13 | 14 | 15 | 16 | Cost | |
| Best-Extension | 12.83 2.56 | 12.83 1.47 | 13.17 1.94 | 15.50 3.62 | 14.17 3.54 | 15.00 2.45 | 15.83 2.32 | 16.17 2.14 | 16.50 1.87 | 16.50 1.22 | 16.83 0.98 | 17.33 1.75 | 16.83 0.75 | 18.17 1.60 | 17.83 1.94 | 47.54% |
| 2 | 3 | 4 | 5 | 6 | 7 | 8 | 9 | Cost | |
| Grok | 58.00 2.28 | 59.83 2.86 | 60.33 2.25 | 60.83 2.32 | 62.50 2.59 | 63.33 1.63 | 63.17 1.94 | 63.00 1.67 | 117.55% |
| 2 | 3 | 4 | 5 | 6 | 7 | 8 | 9 | 10 | 11 | 12 | 13 | 14 | 15 | Cost | |
| GPT | 39.33 3.14 | 42.83 1.72 | 44.00 1.41 | 45.83 3.19 | 47.33 1.75 | 46.33 2.73 | 46.83 3.60 | 46.33 2.73 | 47.83 1.60 | 47.00 2.10 | 46.00 2.37 | 46.67 2.42 | 46.83 2.64 | 46.33 2.42 | 102.71% |
| 2 | 3 | 4 | 5 | 6 | 7 | 8 | 9 | 10 | 11 | 12 | |
| Mistral | 42.67 2.94 | 42.33 2.58 | 43.67 1.97 | 43.33 2.34 | 43.50 0.84 | 44.17 1.33 | 44.33 1.03 | 43.83 0.75 | 43.83 0.41 | 43.50 0.55 | 44.33 1.03 |
| 2 | 3 | 4 | 5 | 6 | 7 | 8 | 9 | 10 | 11 | 12 | Cost | |
| Fuse-Extension | 58.50 2.26 | 61.33 3.56 | 61.17 3.19 | 62.67 2.16 | 60.67 2.58 | 58.00 2.90 | 57.83 3.25 | 57.33 4.84 | 57.33 3.27 | 58.83 1.94 | 59.83 3.06 | 96.24% |
| 2 | 3 | 4 | 5 | 6 | 7 | 8 | 9 | 10 | 11 | Cost | |
| Best-Extension | 59.17 4.40 | 61.33 4.50 | 60.67 3.39 | 61.83 2.64 | 62.83 2.04 | 63.17 2.48 | 61.50 2.81 | 62.17 1.94 | 62.17 2.32 | 62.83 1.72 | 99.14% |
| 2 | 3 | 4 | 5 | 6 | 7 | 8 | 9 | Cost | |
| Grok | 57.83 1.47 | 59.00 1.41 | 60.00 0.63 | 60.17 1.47 | 61.00 1.55 | 61.00 1.41 | 60.50 1.52 | 60.33 1.37 | 115.57% |
| 2 | 3 | 4 | 5 | 6 | 7 | 8 | 9 | 10 | 11 | 12 | 13 | 14 | 15 | 16 | Cost | |
| GPT | 45.50 3.02 | 45.83 2.71 | 47.33 1.21 | 47.00 0.00 | 47.50 1.76 | 47.50 0.84 | 47.33 0.52 | 47.83 1.60 | 47.83 1.33 | 48.50 1.05 | 47.67 0.82 | 47.33 1.03 | 47.17 1.33 | 47.67 1.63 | 47.33 1.63 | 103.85% |
| 2 | 3 | 4 | 5 | 6 | 7 | 8 | 9 | 10 | |
| Mistral | 57.50 2.17 | 58.83 1.72 | 58.00 1.79 | 58.33 2.25 | 58.67 3.39 | 58.33 2.25 | 58.33 2.50 | 58.17 2.64 | 58.67 2.34 |
| 2 | 3 | 4 | 5 | 6 | 7 | 8 | 9 | 10 | 11 | 12 | Cost | |
| Fuse-Extension | 57.67 2.58 | 59.00 1.55 | 59.83 1.33 | 60.83 1.17 | 62.33 2.25 | 64.33 2.07 | 64.33 2.25 | 66.17 2.32 | 65.83 2.23 | 64.50 1.76 | 64.50 2.07 | 99.23% |
| 2 | 3 | 4 | 5 | 6 | 7 | 8 | 9 | 10 | 11 | Cost | |
| Best-Extension | 57.67 2.07 | 59.17 1.83 | 57.83 2.04 | 59.00 2.45 | 60.50 2.35 | 62.50 3.08 | 63.33 3.20 | 64.17 3.06 | 63.67 2.50 | 63.17 3.31 | 99.88% |