Organizations: School of CIT, Technical University of Munich · National Institute of Advanced Industrial Science and Technology (AIST) · The University of Tokyo · Nara Women’s University · Keio University · Carnegie Mellon University
Real-time video commentary generation provides textual descriptions of ongoing events in videos. It supports accessibility and engagement in domains such as sports, esports, and livestreaming. Commentary generation involves two essential decisions: what to say and when to say it. While recent prompting-based approaches using multimodal large language models (MLLMs) have shown strong performance in content generation, they largely ignore the timing aspect. We investigate whether in-context prompting alone can support real-time commentary generation that is both semantically relevant and well-timed. We propose two prompting-based decoding strategies: 1) a fixed-interval approach, and 2) a novel dynamic interval-based decoding approach that adjusts the next prediction timing based on the estimated duration of the previous utterance. Both methods enable pause-aware generation without any fine-tuning. Experiments on Japanese and English datasets of racing and fighting games show that the dynamic interval-based decoding can generate commentary more closely aligned with human utterance timing and content using prompting alone. We release a multilingual benchmark dataset, trained models, and implementations to support future research on real-time video commentary generation.
Figures & tables
Figure 1 : An example of automatically generated commentary for a racing game shown as a subtitle on video.
Figure 2 : Illustration of both Fixed Interval-based and Dynamic Interval-based decoding strategies queried at uniform intervals of t seconds. For Fixed Interval-based strategy, k is fixed as 0.
Video Durations (sec)
Commentary Words (#)
Dataset
avg
min
max
avg
min
max
Race (en)
333.5
254
452
563.6
180
866
Race (jp)
330.8
249
365
681.0
341
1128
Fight (jp)
178.2
171
196
546.1
454
627
Table 1 : Average, Minimum, and Maximum length of Videos in seconds, word count of Reference Commentaries of all 3 datasets
Model
Race (en)
Race (jp)
Fight (jp)
avg
min
max
avg
min
max
avg
min
max
Llava 7b
1358.3
409
3238
870.9
1121.2
184
944.6
0
1868
Qwen 7b
2393.9
424
5309
2989.2
0
6497
677.2
0
3432
GPT-4.1
1241.8
166
3520
1440.0
150
3563
1773.9
306
2615
Reference
563.6
180
866
681.0
341
1128
546.1
454
627
Table 2 : Average words in LLM-generated commentary across all decoding strategies.
Race (English)
Race (Japanese)
Fighting
Alignment
BERTScore
ROUGE
Alignment
BERTScore
ROUGE
Alignment
BERTScore
ROUGE
LLaVA-NeXT-Video-7B-hf
stateless
0.16
0.18
5.1
0.51
0.14
0.1
0.18
0.28
4.9
Feedback
0.65
0.20
10.6
0.20
0.13
0.8
0.21
0.30
1.0
Feedback (ICL)
0.19
0.18
8.1
0.22
0.28
5.6
0.30
0.23
0.7
Realtime
0.25
0.19
6.4
0.19
0.23
0.5
0.18
0.30
0.6
Table 3 : Time Alignment [0 - 1], BERTScore F 1 [-1 - +1], ROUGE-L (0 - 100) of all models across the three datasets and all four decoding strategies.
Figure 3 : Contextual similarities between LLM-generated and Reference Commentaries compute over 10% segments of the whole video.
step size
stateless
feedback
real-time
avg
1
0.166
0.295
0.693
0.462
2
0.163
0.687
0.417
0.421
5
0.172
0.202
0.549
0.368
10
0.169
0.170
0.460
0.3151
Table 4 : Effects of step size on the utterance timing correlation on different decoding strategies
Race (English)
Race (Japanese)
Fighting
KEI
Pause-aware
Coh
Nat
KEI
Pause-aware
Coh
Nat
KEI
Pause-aware
Coh
Nat
LLaVA-NeXT-Video-7B-hf
Feedback (ICL)
2.5
2.0
3.5
1.78
0.0 ∗
0.0 ∗
0.0 ∗
0.0 ∗
0.0 ∗
0.0 ∗
0.0 ∗
0.0 ∗
Realtime
1.85
2.57
2.71
2.07
3.01
3.5
3.89
2.92
2.75
3.15
2.78
2.01
GPT-4.1
Feedback (ICL)
3.00
2.20
3.00
2.60
3.17
2.75
3.00
4.00
3.00
1.93
2.36
3.29
Table 5 : Human Evaluation scores on all 3 datasets across Key Identification Event (KEI), Patience, Conference (Coh), and Naturalness (Nat). A score of 0.0 marked with an asterisk ( ∗ ) indicates cases where the output was either in the wrong language or completely ungrammatical and unintelligible.
Race (En)
Race (Ja)
Fight (Ja)
LLaVA-NeXT-Video-7B-hf
Feedback (ICL)
0.12
0.40
0.17
Realtime
0.046
0.61
0.74
GPT-4.1
Feedback (ICL)
0.55
0.56
0.97
Realtime
0.0
0.0
0.0
Table 6 : Proportion of overlap in commentaries on all 3 datasets on Fixed-Interval and dynamic-interval, both using Feedback-based decoding.
Car Racing (Feedback - ICL)
00:00:00,000 –> 00:00:07,846
Eight racers, Shaun Akiyama, Marquis Morrill, Kurtis Nadeau, Brenton Shull, and Jordon Ordonez, compete in a thrilling race on the Monza Circuit, each driving their unique vehicles.
00:00:18,000 –> 00:00:24,102
Shaun Akiyana leads the pack, followed closely by Marquis Morril and Kurtis Nadau. Brenton Shul and Jordon Ordenez trail further back.
Table 7 : The first four generations by Qwen2.5 using Feedback (ICL) and Realtime decoding method on Race Commentary (en). The examples show that the Feedback method, even with ICL examples, is quite verbose. The Realtime method, on the other hand, generates more concise and equidistant commentaries.
Car Racing (English)
You are a professional commentator for car racing games. You will be provided with a video clip that represents the start of a race. Your task is to generate one sentence of commentary.
1) You should identify the number of players and their names, along with cars.
2) Ignore the background information and refrain from describing the scenery.
3) Initial information about the game without being too verbose.
Car Racing (Japanese)
Table 8 : The prompts used by our methods during the initialization step for all three datasets.
Car Racing (English)
You are a professional commentator for car racing games. You are provided with a video clip from an ongoing car racing game and commentary generated for the game so far.
Previous generated Commentary: {context}
Your task is to compare the given video with the previously generated commentary.
1) Identify if the video has any new development as compared to the already provided commentary.
2) Ignore the background information and refrain from describing the scenery too much.
3) If the state of the game as compared to the provided commentary has not changed, then generate <WAIT>
Table 9 : The prompts used by our methods during inference on all three datasets.
General Instructions
Each score should be evaluated independently (e.g., a commentary may be fluent but still fail to identify key events).
A score of 0 should only be used if the commentary is unreadable or in the wrong language.
When the commentary includes multiple sentences, evaluate based on the overall impression and frequency of issues.
Recommended Evaluation Workflow (Example)
Play or read the generated commentary alongside the video or reference.
Assign a score (0–5) for each of the four criteria.
Table 10 : Evaluation Dimensions and Scoring Criteria for Human Evaluation
We present a low-latency real-time audio game commentary system that generates spoken commentary directly from live gameplay video. In this end-to-end setting, a key bottleneck is accumulated waiting time; conventional pipelines capture frames, generate text, and synthesize speech sequentially for each utterance, and do not request the next generation until speech playback has completed. This strict sequentiality causes long and unnatural silence between utterances. To address this latency bottleneck, our system runs text generation in parallel with speech playback and buffers multiple candidate utterances ahead of time, enabling immediate synthesis at playback boundaries. Experiments on fast-paced game videos show that our parallel design reduces the mean inter-utterance silence from 9.6 seconds to 0.3 seconds compared to sequential baselines. It also improves similarity to professional speaking--silence timing patterns by over 40 %, and a user study with 120 experienced game players confirms significantly improved perceived speaking rhythm. Our demo video is available at: https://youtu.be/pmrRUlvav8M.
Ryota Kawamatsu, Anum Afzal, Yuki Saito +5
The University of Tokyo, Japan · National Institute of Advanced Industrial Science and Technology, Japan · Technical University of Munich, Germany +3
Existing streaming multimodal models process observations incrementally but still follow a turn-based prefill-then-decode pattern, making them non-duplex: new observations cannot naturally enter an active generation stream. Proactive alternatives use micro-turn polling or external response gates, which fragment continuous interaction, decouple response timing from language generation, and complicate KV-cache-friendly serving. We introduce Aero Realtime, a 4B streaming multimodal model with a duplex architecture for realtime generation. Aero Realtime aligns video, audio, and textual output on a shared temporal grid, where each approximately 80-ms audio slot predicts either a lexical token or a silence token. This allows input and output to advance together, enabling one autoregressive objective to learn both when to respond and what to generate. During inference, Aero Realtime appends only the newest multimodal slot, carries forward the previous output state, and reuses the KV cache for efficient incremental execution. We further provide a complete training and serving recipe, including realtime QA construction, slot-aligned supervision, hardware-aware distributed training, and resumable inference. On four NVIDIA A6000 workstation GPUs, Aero Realtime maintains 84-ms median and 173-ms P95 processing lag over 20 minutes of a continuously streamed video, remaining within 200~ms of the source timeline. These results demonstrate the feasibility of fully aligned input-output modeling for duplex, proactive, and hardware-aligned multimodal interaction.
Kaichen Zhang, Wei Huang, Keming Wu +2
1The University of Hong Kong · 2LMMs-Lab · 3Tsinghua University
Real-time commentary for Multiplayer Online Battle Arena (MOBA) esports requires a vision-language model (VLM) to narrate a live match second by second, both fluently and accurately. Existing streaming VLMs sound natural but often miss key events such as kills and objectives. To address this limitation, we use game telemetry, which records exactly when each event occurs, as a supervision signal. We introduce MOBA-VL, a 9B-parameter model trained on this signal with event-localized multi-turn reinforcement learning, which rewards the turns that describe each event. We also collect MOBACast, 860 professional matches (about 460 hours) across three MOBA games with word-level timestamped commentary, and MOBACast-Bench, a benchmark from held-out tournaments. On MOBACast-Bench, MOBA-VL achieves the highest Overall score on full matches (63.25 vs. 55.12 for StreamingVLM) and clips (63.45 vs. 56.22 for DeepSeek-V4.1-Flash). Event-localized credit also raises event recall from 34.5 to 42.1 over supervised fine-tuning. Code and data will be released, and demos are available on an anonymous project page at https://moba-vl.github.io.
Shengyun Zhong, Xinkang Zhao, Ziyuan Chu +1
Zhejiang University, China · Northeastern University, USA