Measuring Collapse and Correction in Homogeneous-Panel LLM Debate
Organizations: Nanyang Technological University
Abstract
Multi-agent large language model (LLM) debate is often evaluated by whether final answers improve, but movement is not necessarily improvement: the same discussion can rescue an initially wrong majority or destroy an initially correct one. Standard final-accuracy evaluations conflate these opposing mechanisms. We introduce an auditable protocol for homogeneous debate on multiple-choice questions (MCQs) that records each run as a transition ledger over collapse, correction, onset, and signed intervention utility. On 6,925 MMLU-Pro debates, the protocol identifies 253 collapses and a parallel correction ledger that changes how interventions should be judged. Replay experiments reveal the central tradeoff: a leave-one-model-out probe-gated freeze prevents 29 collapses but loses 108 corrections under equal weights, so collapse prevention alone can recommend the wrong policy. A compact pre-debate 8-probe screen is a triage signal: its unadjusted family-level association with conditional-collapse risk is high (G=7, Spearman rho=0.893, exact two-sided p=0.0123), but initial-majority accuracy is a close comparator (rho=0.821; family partial rho=0.767, p=0.0877), so we do not treat it as calibrated or capability-adjusted prediction. Round-level traces localize many collapses to the first debate round, where early disagreement can precede both harmful cascades and useful recovery. We release replayable schemas, coders, audits, cost cards, and zero-API rebuild scripts so future model-scaffold rows can be compared under the same denominators and signed utility ledger.
Figures & tables
| Layer | New-row input | Rebuilt outputs and boundary |
|---|---|---|
| Selection | Model/backend, pool id, -probe rows | , channel split, parser/cost audit, and capability comparison; aggregates open, tuned convince-wrong text gated |
| Accounting | Debate JSON Lines (JSONL) with Round and final answers | Initial/final accuracy, transition counts, conditional denominators, and intervals; no LLM judge labels |
| Localization | Round traces or derived Round feature matrix | Onset and Round diagnostics; transcript-derived matrices may be gated, but missingness is explicit |
| Replay | Gate decisions, held-out split, weights | ledger, break-even ratio, and provenance; oracle and collapse-only scores are separated from held-out replay |
| Quantity | Value | Scope |
| Pre-debate screening | ||
| split-half / test–retest / intraclass corr. (ICC) | ; ; – | probe |
| vs. Spearman | ( ) | |
| initial-accuracy comparator; partial after init-acc | ; ( ) | |
| Runtime diagnosis and localization | ||
| disagreement / probe-FR / S/A AUC | – ; – ; – | per question |
| Family | Models | ||
|---|---|---|---|
| DeepSeek | |||
| OpenAI | |||
| Anthropic | |||
| Phi | |||
| Qwen |
| Predictor | Role | Model-row | Family |
|---|---|---|---|
| -probe | pre-debate screen | ||
| Initial-majority accuracy | capability proxy | ||
| Capability pressure ( init acc) | free risk proxy | ||
| Raw debate-revision proxy | revision-quantity baseline | ||
| Social-over-solo lift | narrow conformity proxy | ||
| Final debate accuracy | post-debate descriptive control |
| Predictor | Gemini | Haiku | GPT | Pooled |
|---|---|---|---|---|
| ( ) | ( ) | ( ) | ( ) | |
| Mean probe FR | 0.455 | 0.626 | 0.735 | 0.658 |
| S/A ratio | 0.394 | 0.473 | 0.576 | 0.529 |
| Initial disagreement | 0.765 | 0.550 | 0.858 | 0.573 |
| Combined (FR+Diff+Disagree) | 0.845 | 0.719 | 0.931 | 0.762 |
| Combined (+S/A) | 0.843 | 0.719 | 0.933 | 0.765 |
| Model | Collapses | Round | Round | Round | |
|---|---|---|---|---|---|
| Llama-3.1-8B | 2,003 | 86 | 43 | 21 | 22 |
| Phi-4-mini | 2,161 | 112 | 67 | 17 | 28 |
| Qwen3-4B | 2,161 | 50 | 35 | 9 | 6 |
| Qwen3-8B | 200 | 3 | 3 | 0 | 0 |
| Gemini 3-flash | 200 | 1 | 1 | 0 | 0 |
| GPT-5.4-mini | 200 | 1 | 0 | 0 | 1 |
| Policy | Cohort | Valid. | acc | Prev. | Lost | Net | |
|---|---|---|---|---|---|---|---|
| Probe-gated freeze | OSS | LOMO | pp | ||||
| Round majority change | traces | fixed replay | pp | ||||
| Learned Round stump | traces | strict LOMO | pp |
| Task | Model (mode) | Accuracy (%) | Private peer | |||
|---|---|---|---|---|---|---|
| MMLU-Pro | Qwen3-8B (thinking) | 1,000 | 15/716 | 17/264 | 71.60 / 72.50 | |
| MMLU-Pro | Qwen3-32B (thinking) | 1,000 | 9/754 | 7/212 | 75.40 / 76.30 | |
| MMLU-Pro | R1-Distill-Llama-8B † | 1,000 | 87/472 | 42/344 | 47.20 / 47.20 | |
| MMLU-Pro | gpt-oss-20b (low effort) | 2,000 | 43/1384 | 78/488 | 69.20 / 73.25 | |
| MMLU-Pro | gpt-oss-20b (medium effort) † | 2,000 | 33/1491 | 49/402 | 74.55 / 76.95 | |
| MATH-L5 | R1-Distill-Llama-8B | 1,324 | 72/1037 | 22/80 | 78.32 / 81.42 |
Appendix figures & tables36 assets
Supplementary material from the paper’s appendix.
Appendix
| Term | Meaning | Role in this paper |
|---|---|---|
| Total -probe flip rate | Locked pre-debate screen used for the rank audit | |
| / | Adversarial / corrective probe flip rates | Channel decomposition; not a standalone safety leaderboard |
| Non-adversarial or neutral probe lane where available | Diagnostic channel, not the headline predictor | |
| S/A | Social-over-argument balance or ratio | Construct diagnostic; weak as a per-question oracle |
| FR | Flip rate | Generic answer-revision rate in probes or debates, depending on table context |
| AA / AS / D | Anti-argument / anti-social / default prompt conditions | Channel-separability manipulation |
| Tier | Contents | License/access | Maintenance note |
|---|---|---|---|
| Open code and derived tables | Evaluation scripts, parser, aggregate tables, figures, zero-API rebuild outputs, reuse-card schema, and Round derived-matrix schema | MIT for code; CC-BY 4.0 for tables/schema | Immutable release tag plus DOI archive |
| Open low-risk probes | Initial-answer prompts and non-social weak/moderate/strong probe templates | CC-BY 4.0 subject to upstream benchmark terms | Versioned as alpha-tot-v1.0 |
| Gated probe/trace bundle | Convince-wrong templates, tuned phrasings, full debate transcripts; exact Round matrix until a transcript-free derived matrix is staged | Research-use click-through license with no redistribution or fine-tuning for user-facing persuasion systems | Access log and snapshot metadata retained |
| Metadata and documentation | Datasheet-style card, Croissant/RAI metadata, provenance hashes, parser-failure notes, cost card | Open with the aggregate artifact | Updated only by new versioned release, not in-place edits |
| Mode | Generations on 200 questions | Use and caveat |
|---|---|---|
| Alpha-lite triage | initial answers probe completions | Cheapest screening pass; estimates under one condition/agent and should be followed by richer logging if high. |
| Paper-full screen | initial answer rows probe completions | Full -condition -agent psychometric estimate used for the reported stress rows; more stable, but not always cheaper than a small debate sweep. |
| Direct debate audit | Approximately debate completions | Measures and correction directly on the same 200-question slice, but only after paying for multi-agent traces. |
| Card field | Example value | Source / release note |
|---|---|---|
| Identity and pool | Phi-4-mini, local vLLM, homogeneous 3-agent standard debate; MMLU-Pro high-FR pool | Aggregate row open; model-specific pool selected outcome-blind by probe behavior. |
| Selection screen | , , , S/A over probe rows | Probe templates and aggregate fields open; tuned convince-wrong phrasing is gated. |
| Transition table | debates; initial accuracy , final accuracy , , correction | Full transcripts are gated; aggregate collapse/correction counts are open. |
| Signed utility | LOMO probe-gated freeze at , : prevented, lost, net ( pp) | Policy row is open; the strict matched- DG/DRS matrix is not in the open checkout. |
| Release flags | Aggregates open; full transcripts restricted for public redistribution; exact Round matrix gated unless a transcript-free derived matrix is supplied | State missing or gated matrices explicitly before reuse. |
| Trace family | Rows / states | None count | Rate / note |
|---|---|---|---|
| DeepSeek-v4-flash debate states | 41 init-None debates; 18 final-None debates | ||
| Gemma-4-31B alpha post-answers | |||
| Gemini 3.1 Flash-Lite alpha post-answers | |||
| Qwen3.5-4B alpha post-answers | |||
| Qwen3.6-35B-A3B debate states | 8 answer/final-tag mismatches |
| Trace | Rows | Timestamp / metadata | SHA256 prefix |
|---|---|---|---|
| Gemini 3.1 Flash-Lite debate | 2026-04-26T23:53:42–23:56:02; Gemini API model recorded | 2cfa125250f6 | |
| Qwen3-32B local-vLLM debate | 2026-04-27T09:22:56–11:21:36; local vLLM, qwen3-32b-a7-local | 9a3c0db8c1a7 | |
| Mistral Small 4 OpenRouter debate | 2026-04-27T08:59:28–09:04:00; OpenRouter, Mistral Small 4 | 991e3baae9e8 | |
| Gemma-4-31B alpha lane | older local lane; partial metadata | eb7fcb5a6f27 |
| Predictor of | Spearman | Fisher-z CI | ||
|---|---|---|---|---|
| Probe (8-probe total) ⋆ | n/a | |||
| (right wrong init correct) | ||||
| (wrong right init wrong) | ||||
| Debate FR (revision quantity in debate) | ||||
| Initial accuracy |
| Model | N | Init. Acc. | Debate FR | Probe | Raw Coll. % | Cond. Coll. % [ CI] | S/A |
| Sonnet 4.5 | 120 | 77.5% | 0.705 | 0.428 | 1.67 | 2.15 [0.59, 7.51] | |
| Haiku 4.5 | 500 | 82.2% | 0.512 | 0.491 | 9.20 | 11.19 [8.50, 14.61] | |
| GPT-4o-mini | 498 | 68.5% | 0.458 | 0.337 | 1.61 | 2.35 [1.19, 4.56] | |
| Gemini 2.5 Flash ∗ | 200 | 36.5% | 0.648 | n/a | 0.50 | 1.37 [0.24, 7.36] | |
| GPT-5.4-mini † | 200 | 71.5% | 0.350 | 0.346 | 0.50 | 0.70 [0.12, 3.85] | |
| Gemini 3-flash † | 200 | 85.0% | 0.249 | 0.273 | 0.50 | 0.59 [0.10, 3.26] |
| Model | S/A (default) | 95% CI |
|---|---|---|
| Haiku 4.5 | [ , ] | |
| Sonnet 4.5 | [ , ] | |
| Sonnet 4.6 | [ , ] | |
| Opus 4.5 | [ , ] | |
| Opus 4.6 | [ , ] | |
| GPT-4o-mini | [ , ] |
| Model | S | A | S/A | InitAcc | |
|---|---|---|---|---|---|
| Opus 4.5 | 0.266 | 85.2% | |||
| Opus 4.6 | 0.273 | 85.8% | |||
| Sonnet 4.5 | 0.428 | 83.6% | |||
| Sonnet 4.6 | 0.127 | 83.8% | |||
| Gemini 3-flash | 0.271 | 84.3% | |||
| Gemini 3.1-flash-lite | 0.293 | 81.9% |
| Model | (soc) | (soc) | ||||
|---|---|---|---|---|---|---|
| GPT-4o-mini | 0.337 | 0.266 | 0.113 | 0.289 | 0.348 | 0.267 |
| GPT-5.4 | 0.121 | 0.057 | 0.163 | 0.220 | 0.139 | 0.066 |
| GPT-5.4-mini | 0.346 | 0.273 | 0.115 | 0.393 | 0.354 | 0.273 |
| GPT-5.4-nano | 0.287 | 0.176 | 0.134 | 0.336 | 0.354 | 0.230 |
| Gemini 3-flash | 0.273 | 0.263 | 0.067 | 0.258 | 0.256 | 0.240 |
| Gemini 3.1-flash-lite | 0.310 | 0.296 | 0.099 | 0.281 | 0.327 | 0.309 |
| Model | FR(D) | FR(AS) | FR(AA) | Order | (AA AS) | ||
| Anthropic (5 models, mean ) | |||||||
| Haiku 4.5 | 450 | .178 | .175 | .234 | 0.202 | 9.2e-2 | |
| Sonnet 4.5 | 1,800 | .443 | .355 | .501 | 0.430 | 2.5e-13 ⋆ | |
| Sonnet 4.6 | 1,799 | .136 | .102 | .149 | 0.188 | 2.4e-4 ⋆ | |
| Opus 4.5 | 1,800 | .319 | .234 | .257 | 0.069 | 1.1e-1 | |
| Opus 4.6 | 1,745 | .293 | .240 | .287 | 0.151 | 8.8e-3 ⋆ | |
| Lane | Family | post-None | FR(D) | FR(AS) | FR(AA) | (AA AS) | Status | ||
|---|---|---|---|---|---|---|---|---|---|
| Mistral Small 4 | Mistral | 1,800 | .421 | .066 | .645 | 2.549 | near-complete | ||
| Llama 3.3 70B Instruct | Meta | 1,797 | .470 | .192 | .580 | 1.306 | near-complete | ||
| Llama 4 Scout | Meta | 1,800 | .434 | .426 | .574 | 0.540 | near-complete; parse QC | ||
| Llama 4 Maverick | Meta | 1,794 | .573 | .622 | .680 | 0.234 | near-complete; parse QC | ||
| Grok 4.1 Fast | xAI | 1,799 | .183 | .068 | .184 | 0.497 | near-complete | ||
| HY3 Preview Free | Tencent | 1,785 | .312 | .083 | .402 | 1.037 | near-complete; preview |
| Model | Family | [Wilson 95%] | |
|---|---|---|---|
| sonnet-4.5 | Anthropic | ||
| deepseek-v4-flash § | DeepSeek | ||
| gemini-3-flash | |||
| gemma-4-31b-it-awq | |||
| llama-3.1-8b | Meta | ||
| gpt-4o-mini | OpenAI |
| Statistic | Value |
|---|---|
| Family mean / median aggregation | |
| Family exact permutation | two-sided |
| Family LoFO worst case | |
| Meta+Qwen joint-drop check | ( , descriptive) |
| model-row estimate | |
| Family-clustered bootstrap on model rows | CI |
| Grouping rule | Groups | Spearman | Notes |
|---|---|---|---|
| Canonical family mean | two-sided exact | ||
| Canonical family median | same ranks as mean | ||
| Worst leave-one-family-out | drop Anthropic/DeepSeek/Google/OpenAI tie | ||
| Drop Qwen family | Qwen not required for positive association | ||
| Drop Meta and Qwen | joint upper-tail leverage sensitivity; two-sided exact | ||
| Qwen split by generation | Qwen-3, Qwen-3.5, Qwen-3.6 |
| Held-out family | Observed | Pred. | PI | PI | Rank error | |
|---|---|---|---|---|---|---|
| Anthropic | ||||||
| DeepSeek | ||||||
| Meta | ||||||
| OpenAI | ||||||
| Phi |
| Held-out family | Observed | EIV PI | EIV PI | Covered at |
|---|---|---|---|---|
| Anthropic | yes | |||
| DeepSeek | yes | |||
| yes | ||||
| Meta | yes | |||
| OpenAI | yes | |||
| Phi | yes |
| Runtime diagnostic | Spearman vs. | Interpretation | |
|---|---|---|---|
| Initial disagreement rate | non-unanimous initial panels collapse more often | ||
| Initial unanimity rate | unanimous initial panels are safer in this slice | ||
| Normalized initial-answer entropy | same signal expressed as answer diversity | ||
| At-risk initial disagreement | diversity remains visible after conditioning on initial majority correctness | ||
| Round majority-change rate | strong runtime trajectory score, measured after debate has started |
| Predictor | Partial | ||
|---|---|---|---|
| partial | |||
| partial | |||
| partial |
| Lane | Family | Qids | AS | AA | Init R4 acc. | R3 | R4 | R4 correction | None | |
|---|---|---|---|---|---|---|---|---|---|---|
| Gemini 3 Flash | ( [ ]) | ( [ ]) | ( [ ]) | |||||||
| Gemini 3.1 Flash-Lite | ( [ ]) | ( [ ]) | ( [ ]) | |||||||
| Gemini 3.1 Pro | ( [ ]) | ( [ ]) | ( [ ]) | |||||||
| Mistral Small 4 | Mistral | ( [ ]) | ( [ ]) | ( [ ]) | ||||||
| Grok 4.1 Fast | xAI | ( [ ]) | ( [ ]) | ( [ ]) |
| Run | Init final acc. | Correction | Signed | rows | Scope note | ||
|---|---|---|---|---|---|---|---|
| Gemini 3.1 Flash-Lite | ( [ ]) | ( [ ]) | , | low-collapse boundary | |||
| Mistral Small 4 | ( [ ]) | ( [ ]) | , | dense stress row | |||
| Llama 3.3 70B | ( [ ]) | ( [ ]) | , | provider-routed stress row | |||
| DeepSeek V4 Flash | ( [ ]) | ( [ ]) | , | parser-caveated utility stress row |
| Quantity | Value | Notes |
|---|---|---|
| Alpha rows / questions | / | complete design |
| anti-argument , anti-social , default | ||
| Debates | matched high-FR slice; overlap | |
| Initial final majority accuracy | net positive | |
| Conditional collapse | ( ) | initially-correct majority at risk |
| Conditional correction | ( ) | initially-wrong majority |
| Coefficient | Posterior mean | HDI | ICC |
|---|---|---|---|
| Round majority changed (z) | n/a | ||
| Round agent flips (z) | n/a | ||
| Round unanimous (z) | n/a | ||
| Model-level (z, conditional) | n/a | ||
| Model | |
|---|---|
| Phi- -mini | |
| Llama- - B | |
| Qwen3- B | |
| Qwen3- B |
| Model | Condition | Acc.% | Collapse% | Correct.% | (McNemar) |
|---|---|---|---|---|---|
| Haiku | Default debate | 80.0 | 15.0 | 3.0 | 0.593 |
| Independent debate | 73.0 | 17.0 | 1.0 | ||
| GPT | Default debate | 73.0 | 1.0 | 3.0 | 1.000 |
| Independent debate | 70.0 | 0.0 | 2.0 |
| Model | Method | Acc.% | Col.% | Corr.% | Corr./Col. | Net |
|---|---|---|---|---|---|---|
| Haiku | No debate | 77.5 | 0.0 | 0.0 | n/a | 0 |
| Standard debate | 72.0 | 9.5 | 4.0 | 0.42 | 11 | |
| Shielded | 77.5 | 2.0 | 2.0 | 1.00 | 0 | |
| ACC | 78.0 | 1.5 | 2.0 | 1.33 | +1 | |
| Oracle upper bound | 79.0 | 0.5 | 2.0 | 4.00 | +3 | |
| Phi-4-mini ‡ | No debate | 57.8 | 0.0 | 0.0 | n/a | 0 |
| Model | Policy | Acc.% | Collapse% | Correct.% | Acc. vs Std | |
|---|---|---|---|---|---|---|
| Haiku | Standard | 800 | 81.8 | 1.4 | 2.5 | +0.0 |
| Shielded | 800 | 80.6 | 1.3 | 1.8 | 1.1 | |
| ACC | 800 | 80.6 | 1.4 | 1.9 | 1.1 | |
| Phi-4-mini (800q) | Standard | 800 | 56.9 | 5.8 | 9.1 | +0.0 |
| Shielded | 800 | 57.5 | 5.3 | 6.9 | +0.6 | |
| ACC | 800 | 58.1 | 4.6 | 6.9 | +1.3 |
| Held-out model | Prevented | Lost | Net | acc (pp) | Oracle acc | |
|---|---|---|---|---|---|---|
| Llama-3.1-8B (86 col / 98 corr) | 0.70 | 10/86 | 7/98 | [ ] | ||
| Phi-4-mini (112 col / 174 corr) | 0.70 | 4/112 | 13/174 | [ ] | ||
| Qwen3-4B (50 col / 461 corr) | 0.05 | 15/50 | 87/461 | [ ] | ||
| Qwen3-8B (3 col / 71 corr) | 0.70 | 0/3 | 1/71 | [ ] | ||
| Pooled (frozen LOMO policy) | n/a | 29/251 | 108/804 | n/a |
| Control | Breakeven | |||||
|---|---|---|---|---|---|---|
| Probe-gated freeze | ||||||
| Round majority changed | ||||||
| Learned Round stump |
| Model | exact perm. | max inv. drop | F floor | McNemar | disc. | |
|---|---|---|---|---|---|---|
| Qwen3.5-4B | ||||||
| Qwen3.5-9B | ||||||
| Qwen3.6-27B-FP | ||||||
| Aggregates: median ; Page’s ( ). | ||||||
| Model | F floor | McNemar | Wilson LB | LB | ||
|---|---|---|---|---|---|---|
| Qwen3.5-4B | ✓ | ✓ | ✓ | |||
| Qwen3.5-9B | n/a | n/a | n/a | |||
| Qwen3.6-27B-FP | ✓ | ✓ | ✓ |
| Pair | Holm | Reject | Observed order follows | |
|---|---|---|---|---|
| llama-3.1-8b vs. qwen3.5-9b | yes | initial accuracy | ||
| llama-3.1-8b vs. qwen3.5-4b | yes | initial accuracy | ||
| qwen3.5-9b vs. qwen3.6-27b-fp8 | yes | initial accuracy | ||
| qwen3.5-4b vs. qwen3.6-27b-fp8 | yes | initial accuracy | ||
| qwen3-8b vs. qwen3.5-9b | yes | screen | ||
| qwen3-8b vs. qwen3.5-4b | yes | screen |
| Task | Panel or model | Acc. (%) | Net | Note | |||
|---|---|---|---|---|---|---|---|
| MMLU-Pro | Mistral + Llama-3.3-70B + DeepSeek | 200 | 74.0 / 81.0 | 5/148 | 19/52 | 26 ties | |
| Mistral + DeepSeek + Llama-4-Scout | 200 | 78.5 / 81.5 | 3/157 | 9/43 | 20 ties | ||
| Mistral + DeepSeek + Qwen3-4B | 200 | 71.5 / 71.5 | 7/143 | 7/57 | 28 ties | ||
| GLM-4.6 (reasoning disabled) | 200 | 76.0 / 75.5 | 6/152 | 5/48 | homogeneous | ||
| GSM8K | Mistral | 300 | 96.3 / 97.0 | 1/289 | 3/11 | sparse | |
| Llama-4-Scout | 300 | 96.3 / 95.7 | 2/289 | 0/11 | sparse |
| Task | Model | binary accuracy | expected tie score | ||
|---|---|---|---|---|---|
| MMLU-Pro | GLM-4-9B | ||||
| Gemma-3-4B | |||||
| Granite-4.1-8B | |||||
| Phi-4 | |||||
| Phi-4-mini † | |||||
| Qwen3-0.6B † |
| Task | Model (mode) | Tokens | Accuracy (%) | Private peer | |||
|---|---|---|---|---|---|---|---|
| MMLU-Pro | Qwen3-4B (thinking) | 16,384 | 1,000 | 18/656 | 16/317 | 65.60 / 66.20 | |
| Qwen3-8B (thinking) | 16,384 | 1,000 | 15/716 | 17/264 | 71.60 / 72.50 | ||
| Qwen3-14B (thinking) | 16,384 | 1,000 | 24/737 | 17/241 | 73.70 / 73.80 | ||
| Qwen3-32B (thinking) | 16,384 | 1,000 | 9/754 | 7/212 | 75.40 / 76.30 | ||
| R1-Distill-Llama-8B † | 16,384 | 1,000 | 87/472 | 42/344 | 47.20 / 47.20 | ||
| gpt-oss-20b (low effort) | 16,384 | 2,000 | 43/1384 | 78/488 | 69.20 / 73.25 |