Debate adjudication requires tracking how arguments develop through interaction, yet existing datasets rarely combine fine-grained debate transcripts with professional judgments collected during real competitions under a shared rubric. We introduce a dataset and benchmark for evaluating large language models' understanding of competitive Chinese-language debate at the match, stage, and speaker levels. We organized 182 matches and recruited 120 professional judges, with each match independently adjudicated by three judges using a predefined rubric. After excluding matches with incomplete records, the dataset contains 148 matches, 2,698 stages, and 20,542 exchange units, with manually verified transcripts and segmentation. It preserves original stage scores, match votes, best-debater ballots, and adjudication rationales. We define three tasks: winner-tendency prediction, stage-score prediction, and best-debater prediction. Zero-shot evaluation of multiple large language models yields a highest winner-prediction accuracy of 66.2%, a highest Pearson correlation of 0.250 between model stage scores and mean human ratings, and a highest best-debater prediction accuracy of 56.8%. The dataset and benchmark provide a testbed for studying large language models' understanding of interactive argumentation and their agreement with professional judges.
Figures & tables
Basic info
Text organization
Supervision
Stage-level scoring
Release and privacy
Dataset
Source
Retained debates
Stage
Clash
Judge commentary
Outcome
Stage scores
Best Debater
Competition- time adjudica.
Unified rubric
Rater
Public
Anonymized
ORCHID ( Zhao et al., 2023 )
Re-curated
1,218
✓
×
×
×
×
×
–
–
–
✓
✓
Chen et al. ( Chen et al., 2026b )
Re-curated
94
✓
✓
×
✓
×
×
×
×
LLM
×
×
DEFINED ( Yu et al., 2026 )
Re-curated
108
✓
×
×
×
△
×
×
✓
Expert/ student
×
×
CEDAR ( Lan et al., 2026 )
Re-curated
600
△
×
△
✓
×
△
–
–
–
✓
×
Conch ( Chen et al., 2026a )
Re-curated
3
✓
✓
×
×
×
×
–
–
–
△
△
Table 1: Comparison of selected debate datasets and adjudication benchmarks. ✓ : available; △ : partially available; × : unavailable; –: not applicable.
Figure 1: Example record from the CCDD. Selected stages and illustrative English paraphrases are shown alongside three judges’ scores for the indicated side. Exchanges are linked to their parent stages and do not have independent quality labels. The lower panels illustrate accompanying judge commentary, the match decision, and the best-debater result.
Figure 2: Overview of the Chinese Competitive Debating Dataset and Benchmark.
Predictor
Accuracy
deepseek-v4-flash
.588
deepseek-v4-flash(Reasoning)
.601
deepseek-v4-pro
.615
gemini-3.5-flash-lite
.595
gpt-5.6-sol
.601
gpt-5.6-luna
.561
Table 2: Winner-tendency prediction results.
Predictor
MSE
Pearson r
Spearman ρ
Agreement
Non-tie Agreement
deepseek-v4-flash
2.72
.250
.253
.402
.638
deepseek-v4-flash (Reasoning)
2.54
.241
.241
.403
.602
deepseek-v4-pro
2.13
.231
.240
.434
.596
gemini-3.5-flash-lite
1.91
.184
.181
.391
.529
gpt-5.6-luna
2.52
.232
.237
.360
.579
Random baseline
9.69
≈0
≈0
.239
.451
Table 3: Stage-score prediction results.
Predictor
Best-Debater Accuracy
Vote MAE
Vote-Share r
Vote-Share ρ
Predicted-Winner Third-Speaker Accuracy
deepseek-v4-flash
.432
.125
.481
.488
.457
deepseek-v4-flash (Reasoning)
.438
.129
.456
.470
.461
deepseek-v4-pro
.554
.118
.549
.534
.491
gemini-3.5-flash-lite
.345
.136
.393
.425
.474
gpt-5.6-sol
.568
.117
.561
.560
.439
gpt-5.6-luna
.432
.122
.508
.524
.422
Table 4: Best-debater and vote-distribution prediction results.
Appendix figures & tables5 assets
Supplementary material from the paper’s appendix.
Appendix
Figure 3: A typical format of Chinese-language competitive debate.
Task Performance
Battlefield Judgment
Degree of Advancement
Suggested Score
Perfectly completed
Correct and important battlefield
Decisive advancement
10
Perfectly completed
Correct and important battlefield
Major advancement
9
Well completed
Correct and important battlefield
Substantial advancement
8
Well completed
Correct and important battlefield
Effective advancement
7
Well completed
Correct and important battlefield
An attempt to advance
6
Well completed
Correct but secondary battlefield
Effective advancement
6
Appendix
Table 5: Stage-level evaluation rubric used by professional judges.
Predictor
Accuracy
deepseek-v4-flash
.588
deepseek-v4-flash(Reasoning)
.601
deepseek-v4-pro
.615
gemini-3.5-flash-lite
.595
gpt-5.6-sol
.601
gpt-5.6-luna
.561
Appendix
Table 6: Winner-tendency prediction results.
Predictor
MSE
Pearson r
Spearman ρ
Agreement
Non-tie Agreement
deepseek-v4-flash
2.72
.250
.253
.402
.638
deepseek-v4-flash (Reasoning)
2.54
.241
.241
.403
.602
deepseek-v4-pro
2.13
.231
.240
.434
.596
gemini-3.5-flash-lite
1.91
.184
.181
.391
.529
gpt-5.6-luna
2.52
.232
.237
.360
.579
Random baseline
9.69
≈0
≈0
.239
.451
Appendix
Table 7: Stage-score prediction results.
Predictor
Best-Debater Accuracy
Vote MAE
Vote-Share r
Vote-Share ρ
Predicted-Winner Third-Speaker Accuracy
deepseek-v4-flash
.432
.125
.481
.488
.457
deepseek-v4-flash (Reasoning)
.439
.129
.456
.470
.461
deepseek-v4-pro
.554
.118
.549
.534
.491
gemini-3.5-flash-lite
.345
.136
.393
.425
.474
gpt-5.6-sol
.568
.117
.561
.560
.439
gpt-5.6-luna
.432
.122
.508
.524
.422
Appendix
Table 8: Best-debater and vote-distribution prediction results. Results marked with ∗ are non-comparable references and are not used for direct model comparison.
LLM debate is usually evaluated by final answers, yet transcripts reveal whether later turns develop new arguments or return to earlier claims in new wording. We study this process with \textit{prior-argument similarity}, which compares extracted argument units with earlier units in the same debate. In controlled eight-turn debates over 71 motions, six languages, and four model agents, Chinese is the only tested language with a consistently positive gap relative to English across three multilingual embedding models. The gap persists across agents, turn positions, regression adjustment, metric variants, extraction-length controls, a second-extractor subset, and cross-encoder tail rescoring. Manual calibration shows weak item-level alignment but a high-similarity tail enriched for substantive repetition. A diversity-aware prompt lowers \textit{prior-argument similarity} across languages, yet does not significantly narrow the Chinese--English gap. Multilingual debate evaluation should therefore measure argumentative development over time and report both average and gap terms.
When should a language model answer directly, sample and vote, or engage in multi-agent debate? Recent work shows voting often explains much of the gain attributed to debate, while selective-debate systems activate deliberation only on uncertain examples. We ask: under a matched ceiling on generated tokens (960 per example), how much per-example routing headroom exists, and how much is recoverable from cheap pre-deliberation signals? We evaluate greedy decoding, three-sample voting, and a two-agent critique-revise debate on MuSiQue and GSM8K using Llama 3.1 8B Instruct and Ministral 3 8B Instruct. On MuSiQue, an oracle selecting the correct protocol per example gains +14.0 and +13.7 pp over the best fixed one. The best fixed protocol is model- and dataset-dependent: each (model, dataset) cell has a different winner. This headroom is hard to recover from cheap ex-ante signals. A vote-entropy threshold is the only controller that directionally beats the best fixed protocol on both models (+1.3 and +1.7 pp), though individual paired-bootstrap CIs include zero. A joint analysis (meta-analysis +1.6 pp, p=0.125; Bayesian P(both>0)=0.59) is directionally consistent but not significant. Learned controllers (LR, GBT) do not outperform the threshold. The key finding is structural: vote entropy predicts where debate is safe, not where debate is needed. High entropy sharply reduces debate backfire, but 66% of debate-helpful examples (31/47) occur when voting is unanimous but wrong. A single-prompt self-critique probe on Llama flips the answer in 127/127 unanimous cases, yielding zero mutual information with the debate-helpful label; we cannot rule out a prompt-compliance artifact, but either interpretation disqualifies the probe as a router. Recovering the remaining headroom requires behavioral probes that avoid format-compliance confounds at the 8B scale.
We extend the DS@GT ARC working-note submission to the Touché 2025 Retrieval-Augmented Debate task. The task has two subtasks: generating the next utterance in a simulated debate, and evaluating debate responses according to the Gricean maxims of Quantity, Quality, Relation, and Manner. The DS@GT ARC submission consisted of six leading LLMs from three providers through a retrieval-augmented prompting pipeline. We summarize the results from the working paper and explore whether multi-LLM evaluator agreement is a reliable proxy for official evaluation performance. The analysis shows that frontier LLM systems are strong response generators, and as evaluators they agree strongly within model families. However this consensus does not reliably track the official evaluation target, with the largest gap on the Quality maxim. The accompanying source code for this paper is located at https://github.com/dsgt-arc/touche-2025-rad and https://github.com/dsgt-arc/touche-2025-rad-analysis.
Anthony Miyaguchi, Conor Johnston
Georgia Institute of Technology, North Ave NW, Atlanta, GA 30332, USA