Organizations: School of CIT, Technical University of Munich · National Institute of Advanced Industrial Science and Technology (AIST) · The University of Tokyo · Nara Women’s University · Keio University · Carnegie Mellon University
Real-time video commentary generation provides textual descriptions of ongoing events in videos. It supports accessibility and engagement in domains such as sports, esports, and livestreaming. Commentary generation involves two essential decisions: what to say and when to say it. While recent prompting-based approaches using multimodal large language models (MLLMs) have shown strong performance in content generation, they largely ignore the timing aspect. We investigate whether in-context prompting alone can support real-time commentary generation that is both semantically relevant and well-timed. We propose two prompting-based decoding strategies: 1) a fixed-interval approach, and 2) a novel dynamic interval-based decoding approach that adjusts the next prediction timing based on the estimated duration of the previous utterance. Both methods enable pause-aware generation without any fine-tuning. Experiments on Japanese and English datasets of racing and fighting games show that the dynamic interval-based decoding can generate commentary more closely aligned with human utterance timing and content using prompting alone. We release a multilingual benchmark dataset, trained models, and implementations to support future research on real-time video commentary generation.
Figures & tables
Figure 1 : An example of automatically generated commentary for a racing game shown as a subtitle on video.
Figure 2 : Illustration of both Fixed Interval-based and Dynamic Interval-based decoding strategies queried at uniform intervals of t seconds. For Fixed Interval-based strategy, k is fixed as 0.
Video Durations (sec)
Commentary Words (#)
Dataset
avg
min
max
avg
min
max
Race (en)
333.5
254
452
563.6
180
866
Race (jp)
330.8
249
365
681.0
341
1128
Fight (jp)
178.2
171
196
546.1
454
627
Table 1 : Average, Minimum, and Maximum length of Videos in seconds, word count of Reference Commentaries of all 3 datasets
Model
Race (en)
Race (jp)
Fight (jp)
avg
min
max
avg
min
max
avg
min
max
Llava 7b
1358.3
409
3238
870.9
1121.2
184
944.6
0
1868
Qwen 7b
2393.9
424
5309
2989.2
0
6497
677.2
0
3432
GPT-4.1
1241.8
166
3520
1440.0
150
3563
1773.9
306
2615
Reference
563.6
180
866
681.0
341
1128
546.1
454
627
Table 2 : Average words in LLM-generated commentary across all decoding strategies.
Race (English)
Race (Japanese)
Fighting
Alignment
BERTScore
ROUGE
Alignment
BERTScore
ROUGE
Alignment
BERTScore
ROUGE
LLaVA-NeXT-Video-7B-hf
stateless
0.16
0.18
5.1
0.51
0.14
0.1
0.18
0.28
4.9
Feedback
0.65
0.20
10.6
0.20
0.13
0.8
0.21
0.30
1.0
Feedback (ICL)
0.19
0.18
8.1
0.22
0.28
5.6
0.30
0.23
0.7
Realtime
0.25
0.19
6.4
0.19
0.23
0.5
0.18
0.30
0.6
Table 3 : Time Alignment [0 - 1], BERTScore F 1 [-1 - +1], ROUGE-L (0 - 100) of all models across the three datasets and all four decoding strategies.
Figure 3 : Contextual similarities between LLM-generated and Reference Commentaries compute over 10% segments of the whole video.
step size
stateless
feedback
real-time
avg
1
0.166
0.295
0.693
0.462
2
0.163
0.687
0.417
0.421
5
0.172
0.202
0.549
0.368
10
0.169
0.170
0.460
0.3151
Table 4 : Effects of step size on the utterance timing correlation on different decoding strategies
Race (English)
Race (Japanese)
Fighting
KEI
Pause-aware
Coh
Nat
KEI
Pause-aware
Coh
Nat
KEI
Pause-aware
Coh
Nat
LLaVA-NeXT-Video-7B-hf
Feedback (ICL)
2.5
2.0
3.5
1.78
0.0 ∗
0.0 ∗
0.0 ∗
0.0 ∗
0.0 ∗
0.0 ∗
0.0 ∗
0.0 ∗
Realtime
1.85
2.57
2.71
2.07
3.01
3.5
3.89
2.92
2.75
3.15
2.78
2.01
GPT-4.1
Feedback (ICL)
3.00
2.20
3.00
2.60
3.17
2.75
3.00
4.00
3.00
1.93
2.36
3.29
Table 5 : Human Evaluation scores on all 3 datasets across Key Identification Event (KEI), Patience, Conference (Coh), and Naturalness (Nat). A score of 0.0 marked with an asterisk ( ∗ ) indicates cases where the output was either in the wrong language or completely ungrammatical and unintelligible.
Race (En)
Race (Ja)
Fight (Ja)
LLaVA-NeXT-Video-7B-hf
Feedback (ICL)
0.12
0.40
0.17
Realtime
0.046
0.61
0.74
GPT-4.1
Feedback (ICL)
0.55
0.56
0.97
Realtime
0.0
0.0
0.0
Table 6 : Proportion of overlap in commentaries on all 3 datasets on Fixed-Interval and dynamic-interval, both using Feedback-based decoding.
Car Racing (Feedback - ICL)
00:00:00,000 –> 00:00:07,846
Eight racers, Shaun Akiyama, Marquis Morrill, Kurtis Nadeau, Brenton Shull, and Jordon Ordonez, compete in a thrilling race on the Monza Circuit, each driving their unique vehicles.
00:00:18,000 –> 00:00:24,102
Shaun Akiyana leads the pack, followed closely by Marquis Morril and Kurtis Nadau. Brenton Shul and Jordon Ordenez trail further back.
Table 7 : The first four generations by Qwen2.5 using Feedback (ICL) and Realtime decoding method on Race Commentary (en). The examples show that the Feedback method, even with ICL examples, is quite verbose. The Realtime method, on the other hand, generates more concise and equidistant commentaries.
Car Racing (English)
You are a professional commentator for car racing games. You will be provided with a video clip that represents the start of a race. Your task is to generate one sentence of commentary.
1) You should identify the number of players and their names, along with cars.
2) Ignore the background information and refrain from describing the scenery.
3) Initial information about the game without being too verbose.
Car Racing (Japanese)
Table 8 : The prompts used by our methods during the initialization step for all three datasets.
Car Racing (English)
You are a professional commentator for car racing games. You are provided with a video clip from an ongoing car racing game and commentary generated for the game so far.
Previous generated Commentary: {context}
Your task is to compare the given video with the previously generated commentary.
1) Identify if the video has any new development as compared to the already provided commentary.
2) Ignore the background information and refrain from describing the scenery too much.
3) If the state of the game as compared to the provided commentary has not changed, then generate <WAIT>
Table 9 : The prompts used by our methods during inference on all three datasets.
General Instructions
Each score should be evaluated independently (e.g., a commentary may be fluent but still fail to identify key events).
A score of 0 should only be used if the commentary is unreadable or in the wrong language.
When the commentary includes multiple sentences, evaluate based on the overall impression and frequency of issues.
Recommended Evaluation Workflow (Example)
Play or read the generated commentary alongside the video or reference.
Assign a score (0–5) for each of the four criteria.
Table 10 : Evaluation Dimensions and Scoring Criteria for Human Evaluation