Game commentary is an open-ended generation task requiring multimodal perception, strategic reasoning, and contextual knowledge. Existing AI-Generated Game Commentary (AI-GGC) studies remain fragmented across games, modalities, and evaluation protocols, while overlap-based or holistic evaluators fail to capture the functional heterogeneity of commentary. We introduce \textsc{GameCommBench}, a unified benchmark spanning board games, sports, and esports, with commentary aligned to heterogeneous game contexts and annotated by commentary type. We further propose Type-Aware Commentary Evaluation (TACE), a structured framework for evaluating different types of commentary. We then validate TACE for reliability and human agreement, and use it to benchmark representative AI commentators. Results reveal non-uniform capability profiles, with live observation and strategic analysis emerging as major bottlenecks. Together, \textsc{GameCommBench} and TACE provide a diagnostic foundation for comparable and interpretable AI-GGC evaluation.
Figures & tables
Figure 1: Existing AI-GGC research is fragmented across domains and modalities, and lacks effective evaluation schemes. We address these issues with GameCommBench , a unified cross-domain benchmark, and TACE, a commentary evaluation framework designed for descriptive, analytical, and background commentary.
Figure 2: Overview of GameCommBench . The figure summarizes benchmark coverage, representative state-commentary examples, and commentary-type distributions across games. The central wheel reports the proportions of descriptive, analytical, and background commentary for each game. Scatter plots visualize classified human commentary embeddings, with red, blue, and green denoting descriptive, analytical, and background commentary.
Figure 3: Overview of TACE. TACE evaluates descriptive, analytical, and background commentary through three type-specific pipelines. The resulting metrics assess the description correctness and completeness, reasoning faithfulness and depth, and background factuality and relevance.
Table 1: Validation of TACE. We report run-to-run reliability and human validation for LLM-mediated components.
Figure 4: Main benchmark results on GameCommBench . The top table reports aggregated model scores for the six TACE metrics, with the average row summarizing metric-wise performance across evaluated models. The bottom heatmaps show game-wise capability profiles across models and metrics, revealing cross-game and type-specific performance patterns.
Table 2: Effects of model scale and reasoning mode. Left: comparison of open-weight models with different sizes. Right: comparison between minimal-reasoning and thinking settings on board games.
Figure 5: Detailed data processing pipeline. Raw game data are collected and aligned, cleaned through quality control and temporal repair, decomposed into commentary types, and packaged into a unified format.
Appendix figures & tables4 assets
Supplementary material from the paper’s appendix.
Appendix
Game
Primary alignment
Processing summary
Football
Caption/event clip to transcript
Segment and classify transcript, filter against caption, fuse caption and retained commentary, reclassify, then normalize IDs and labels.
Basketball
NSVA caption/action clip to transcript
Gate caption-commentary relevance, segment/classify, filter, fuse, select high-quality merged outputs, then normalize with NBA metadata and action labels.
Chess
PGN mainline ply
Parse GameKnot HTML into PGN comments; preserve move-attached commentary with only lightweight text normalization.
Go
SGF mainline ply
Crawl FoxWQ JSON, extract SGF, discard empty or irrelevant comments, and normalize turn-aligned records.
League of Legends
Subtitle window plus API window
Align subtitle segments to clips, aggregate Riot/LoL Esports API windows, and retain significant team/player state changes.
Appendix
Table 3: Integrated data-processing workflows for the five game types.
Game
Input supplied to model
Decoding policy
State construction
Football
Timestamped video frames, match background, clip duration
temperature 0; thinking disabled
4 FPS visual sampling with original frame resolution.
Basketball
Timestamped video frames, match background, clip duration
temperature 0; thinking disabled
4 FPS visual sampling with original frame resolution.
ASCII board, SGF and standard coordinate, recent moves, background
temperature 0; thinking disabled
SGF mainline parsing with board rendering before the current move.
League of Legends
Video input, match background, API transition context
temperature 0; thinking disabled
Clip-aligned visual input plus team/player state transitions.
Appendix
Table 4: Candidate commentary generation settings used by the provided scripts.
Game
Accuracy
Chess
94.0%
Go
93.0%
Football
98.0%
Basketball
98.0%
LoL
96.0%
Overall
95.8%
Appendix
Table 5: Human validation accuracy of LLM-based type-aware decomposition across games.
Type
Metric
Spearman ρ
Analytical
Faithfulness
0.77
Reasoning Depth
0.94
Appendix
Table 6: Cross-judge consistency for analytical metrics. We compare model rankings produced by the main GPT-5.4-mini-based judge and an alternative Gemini-3-Flash-based judge on a held-out subset.