Debate adjudication requires tracking how arguments develop through interaction, yet existing datasets rarely combine fine-grained debate transcripts with professional judgments collected during real competitions under a shared rubric. We introduce a dataset and benchmark for evaluating large language models' understanding of competitive Chinese-language debate at the match, stage, and speaker levels. We organized 182 matches and recruited 120 professional judges, with each match independently adjudicated by three judges using a predefined rubric. After excluding matches with incomplete records, the dataset contains 148 matches, 2,698 stages, and 20,542 exchange units, with manually verified transcripts and segmentation. It preserves original stage scores, match votes, best-debater ballots, and adjudication rationales. We define three tasks: winner-tendency prediction, stage-score prediction, and best-debater prediction. Zero-shot evaluation of multiple large language models yields a highest winner-prediction accuracy of 66.2%, a highest Pearson correlation of 0.250 between model stage scores and mean human ratings, and a highest best-debater prediction accuracy of 56.8%. The dataset and benchmark provide a testbed for studying large language models' understanding of interactive argumentation and their agreement with professional judges.
Figures & tables
Basic info
Text organization
Supervision
Stage-level scoring
Release and privacy
Dataset
Source
Retained debates
Stage
Clash
Judge commentary
Outcome
Stage scores
Best Debater
Competition- time adjudica.
Unified rubric
Rater
Public
Anonymized
ORCHID ( Zhao et al., 2023 )
Re-curated
1,218
✓
×
×
×
×
×
–
–
–
✓
✓
Chen et al. ( Chen et al., 2026b )
Re-curated
94
✓
✓
×
✓
×
×
×
×
LLM
×
×
DEFINED ( Yu et al., 2026 )
Re-curated
108
✓
×
×
×
△
×
×
✓
Expert/ student
×
×
CEDAR ( Lan et al., 2026 )
Re-curated
600
△
×
△
✓
×
△
–
–
–
✓
×
Conch ( Chen et al., 2026a )
Re-curated
3
✓
✓
×
×
×
×
–
–
–
△
△
Table 1: Comparison of selected debate datasets and adjudication benchmarks. ✓ : available; △ : partially available; × : unavailable; –: not applicable.
Figure 1: Example record from the CCDD. Selected stages and illustrative English paraphrases are shown alongside three judges’ scores for the indicated side. Exchanges are linked to their parent stages and do not have independent quality labels. The lower panels illustrate accompanying judge commentary, the match decision, and the best-debater result.
Figure 2: Overview of the Chinese Competitive Debating Dataset and Benchmark.
Predictor
Accuracy
deepseek-v4-flash
.588
deepseek-v4-flash(Reasoning)
.601
deepseek-v4-pro
.615
gemini-3.5-flash-lite
.595
gpt-5.6-sol
.601
gpt-5.6-luna
.561
Table 2: Winner-tendency prediction results.
Predictor
MSE
Pearson r
Spearman ρ
Agreement
Non-tie Agreement
deepseek-v4-flash
2.72
.250
.253
.402
.638
deepseek-v4-flash (Reasoning)
2.54
.241
.241
.403
.602
deepseek-v4-pro
2.13
.231
.240
.434
.596
gemini-3.5-flash-lite
1.91
.184
.181
.391
.529
gpt-5.6-luna
2.52
.232
.237
.360
.579
Random baseline
9.69
≈0
≈0
.239
.451
Table 3: Stage-score prediction results.
Predictor
Best-Debater Accuracy
Vote MAE
Vote-Share r
Vote-Share ρ
Predicted-Winner Third-Speaker Accuracy
deepseek-v4-flash
.432
.125
.481
.488
.457
deepseek-v4-flash (Reasoning)
.438
.129
.456
.470
.461
deepseek-v4-pro
.554
.118
.549
.534
.491
gemini-3.5-flash-lite
.345
.136
.393
.425
.474
gpt-5.6-sol
.568
.117
.561
.560
.439
gpt-5.6-luna
.432
.122
.508
.524
.422
Table 4: Best-debater and vote-distribution prediction results.
Appendix figures & tables5 assets
Supplementary material from the paper’s appendix.
Appendix
Figure 3: A typical format of Chinese-language competitive debate.
Task Performance
Battlefield Judgment
Degree of Advancement
Suggested Score
Perfectly completed
Correct and important battlefield
Decisive advancement
10
Perfectly completed
Correct and important battlefield
Major advancement
9
Well completed
Correct and important battlefield
Substantial advancement
8
Well completed
Correct and important battlefield
Effective advancement
7
Well completed
Correct and important battlefield
An attempt to advance
6
Well completed
Correct but secondary battlefield
Effective advancement
6
Appendix
Table 5: Stage-level evaluation rubric used by professional judges.
Predictor
Accuracy
deepseek-v4-flash
.588
deepseek-v4-flash(Reasoning)
.601
deepseek-v4-pro
.615
gemini-3.5-flash-lite
.595
gpt-5.6-sol
.601
gpt-5.6-luna
.561
Appendix
Table 6: Winner-tendency prediction results.
Predictor
MSE
Pearson r
Spearman ρ
Agreement
Non-tie Agreement
deepseek-v4-flash
2.72
.250
.253
.402
.638
deepseek-v4-flash (Reasoning)
2.54
.241
.241
.403
.602
deepseek-v4-pro
2.13
.231
.240
.434
.596
gemini-3.5-flash-lite
1.91
.184
.181
.391
.529
gpt-5.6-luna
2.52
.232
.237
.360
.579
Random baseline
9.69
≈0
≈0
.239
.451
Appendix
Table 7: Stage-score prediction results.
Predictor
Best-Debater Accuracy
Vote MAE
Vote-Share r
Vote-Share ρ
Predicted-Winner Third-Speaker Accuracy
deepseek-v4-flash
.432
.125
.481
.488
.457
deepseek-v4-flash (Reasoning)
.439
.129
.456
.470
.461
deepseek-v4-pro
.554
.118
.549
.534
.491
gemini-3.5-flash-lite
.345
.136
.393
.425
.474
gpt-5.6-sol
.568
.117
.561
.560
.439
gpt-5.6-luna
.432
.122
.508
.524
.422
Appendix
Table 8: Best-debater and vote-distribution prediction results. Results marked with ∗ are non-comparable references and are not used for direct model comparison.