Game commentary is an open-ended generation task requiring multimodal perception, strategic reasoning, and contextual knowledge. Existing AI-Generated Game Commentary (AI-GGC) studies remain fragmented across games, modalities, and evaluation protocols, while overlap-based or holistic evaluators fail to capture the functional heterogeneity of commentary. We introduce \textsc{GameCommBench}, a unified benchmark spanning board games, sports, and esports, with commentary aligned to heterogeneous game contexts and annotated by commentary type. We further propose Type-Aware Commentary Evaluation (TACE), a structured framework for evaluating different types of commentary. We then validate TACE for reliability and human agreement, and use it to benchmark representative AI commentators. Results reveal non-uniform capability profiles, with live observation and strategic analysis emerging as major bottlenecks. Together, \textsc{GameCommBench} and TACE provide a diagnostic foundation for comparable and interpretable AI-GGC evaluation.
Figures & tables
Figure 1: Existing AI-GGC research is fragmented across domains and modalities, and lacks effective evaluation schemes. We address these issues with GameCommBench , a unified cross-domain benchmark, and TACE, a commentary evaluation framework designed for descriptive, analytical, and background commentary.
Figure 2: Overview of GameCommBench . The figure summarizes benchmark coverage, representative state-commentary examples, and commentary-type distributions across games. The central wheel reports the proportions of descriptive, analytical, and background commentary for each game. Scatter plots visualize classified human commentary embeddings, with red, blue, and green denoting descriptive, analytical, and background commentary.
Figure 3: Overview of TACE. TACE evaluates descriptive, analytical, and background commentary through three type-specific pipelines. The resulting metrics assess the description correctness and completeness, reasoning faithfulness and depth, and background factuality and relevance.
Table 1: Validation of TACE. We report run-to-run reliability and human validation for LLM-mediated components.
Figure 4: Main benchmark results on GameCommBench . The top table reports aggregated model scores for the six TACE metrics, with the average row summarizing metric-wise performance across evaluated models. The bottom heatmaps show game-wise capability profiles across models and metrics, revealing cross-game and type-specific performance patterns.
Table 2: Effects of model scale and reasoning mode. Left: comparison of open-weight models with different sizes. Right: comparison between minimal-reasoning and thinking settings on board games.
Figure 5: Detailed data processing pipeline. Raw game data are collected and aligned, cleaned through quality control and temporal repair, decomposed into commentary types, and packaged into a unified format.
Appendix figures & tables4 assets
Supplementary material from the paper’s appendix.
Appendix
Game
Primary alignment
Processing summary
Football
Caption/event clip to transcript
Segment and classify transcript, filter against caption, fuse caption and retained commentary, reclassify, then normalize IDs and labels.
Basketball
NSVA caption/action clip to transcript
Gate caption-commentary relevance, segment/classify, filter, fuse, select high-quality merged outputs, then normalize with NBA metadata and action labels.
Chess
PGN mainline ply
Parse GameKnot HTML into PGN comments; preserve move-attached commentary with only lightweight text normalization.
Go
SGF mainline ply
Crawl FoxWQ JSON, extract SGF, discard empty or irrelevant comments, and normalize turn-aligned records.
League of Legends
Subtitle window plus API window
Align subtitle segments to clips, aggregate Riot/LoL Esports API windows, and retain significant team/player state changes.
Appendix
Table 3: Integrated data-processing workflows for the five game types.
Game
Input supplied to model
Decoding policy
State construction
Football
Timestamped video frames, match background, clip duration
temperature 0; thinking disabled
4 FPS visual sampling with original frame resolution.
Basketball
Timestamped video frames, match background, clip duration
temperature 0; thinking disabled
4 FPS visual sampling with original frame resolution.
ASCII board, SGF and standard coordinate, recent moves, background
temperature 0; thinking disabled
SGF mainline parsing with board rendering before the current move.
League of Legends
Video input, match background, API transition context
temperature 0; thinking disabled
Clip-aligned visual input plus team/player state transitions.
Appendix
Table 4: Candidate commentary generation settings used by the provided scripts.
Game
Accuracy
Chess
94.0%
Go
93.0%
Football
98.0%
Basketball
98.0%
LoL
96.0%
Overall
95.8%
Appendix
Table 5: Human validation accuracy of LLM-based type-aware decomposition across games.
Type
Metric
Spearman ρ
Analytical
Faithfulness
0.77
Reasoning Depth
0.94
Appendix
Table 6: Cross-judge consistency for analytical metrics. We compare model rankings produced by the main GPT-5.4-mini-based judge and an alternative Gemini-3-Flash-based judge on a held-out subset.
The advent of artificial intelligence has propelled AI-Generated Game Commentary (AI-GGC) into a rapidly expanding research area, offering advantages such as scalable availability and personalized narration. However, existing studies remain fragmented, and a systematic survey that unifies prior efforts is still lacking. To bridge this gap, our survey introduces a unified framework that systematically organizes the AI-GGC landscape. We present a novel taxonomy focused on three core commentator capabilities: Live Observation, Strategic Analysis, and Historical Recall, and further categorize commentary into three corresponding types: Descriptive Commentary, Analytical Commentary, and Background Commentary. Building on this structure, we provide an in-depth review of methods, datasets, and evaluation metrics, analyzing their strengths and limitations. Finally, we highlight key challenges and point out promising directions for future research in AI-GGC.
Qirui Zheng, Xingbo Wang, Keyuan Cheng +5
Peking University · King Abdullah University of Science and Technology
Real-time video commentary generation provides textual descriptions of ongoing events in videos. It supports accessibility and engagement in domains such as sports, esports, and livestreaming. Commentary generation involves two essential decisions: what to say and when to say it. While recent prompting-based approaches using multimodal large language models (MLLMs) have shown strong performance in content generation, they largely ignore the timing aspect. We investigate whether in-context prompting alone can support real-time commentary generation that is both semantically relevant and well-timed. We propose two prompting-based decoding strategies: 1) a fixed-interval approach, and 2) a novel dynamic interval-based decoding approach that adjusts the next prediction timing based on the estimated duration of the previous utterance. Both methods enable pause-aware generation without any fine-tuning. Experiments on Japanese and English datasets of racing and fighting games show that the dynamic interval-based decoding can generate commentary more closely aligned with human utterance timing and content using prompting alone. We release a multilingual benchmark dataset, trained models, and implementations to support future research on real-time video commentary generation.
Anum Afzal, Yuki Saito, Hiroya Takamura +5
School of CIT, Technical University of Munich · National Institute of Advanced Industrial Science and Technology (AIST) · The University of Tokyo +3
Game generation is an emerging application of coding agents, requiring models to transform natural-language specifications into playable interactive systems. Unlike traditional coding tasks, game generation takes place within a game engine, where scripts, scenes, assets, rendering, and runtime interactions must jointly produce coherent gameplay. We formalize end-to-end game generation as the problem of producing a complete game artifact that realizes a specification through observable player-game interaction in a target environment. We argue that evaluating this setting requires three desiderata: Engine Grounding, Artifact Completeness, and Interactive Verification. We propose an interaction-grounded evaluation framework that assesses executable gameplay through replayed demonstrations and rubric-guided multimodal judging. We instantiate this framework as GameCraft-Bench, a benchmark comprising 140 Godot tasks across 15 game families. Evaluations of frontier coding agents show that end-to-end game generation remains highly challenging: the strongest agent achieves only 41.46%, and most agents score below 40%. Further analysis reveals that while agents often implement recognizable mechanics, they struggle to deliver complete games with sufficient content, functional visual feedback, and coherent presentation. See https://tongxuluo.github.io/gamecraft-bench-website for demos, code, and data.
Tongxu Luo, Rongsheng Wang, Jiaxi Bi +22
1The Chinese University of Hong Kong, Shenzhen · 2Shenzhen Loop Area Institute · 4USTB +5