Real-time commentary for Multiplayer Online Battle Arena (MOBA) esports requires a vision-language model (VLM) to narrate a live match second by second, both fluently and accurately. Existing streaming VLMs sound natural but often miss key events such as kills and objectives. To address this limitation, we use game telemetry, which records exactly when each event occurs, as a supervision signal. We introduce MOBA-VL, a 9B-parameter model trained on this signal with event-localized multi-turn reinforcement learning, which rewards the turns that describe each event. We also collect MOBACast, 860 professional matches (about 460 hours) across three MOBA games with word-level timestamped commentary, and MOBACast-Bench, a benchmark from held-out tournaments. On MOBACast-Bench, MOBA-VL achieves the highest Overall score on full matches (63.25 vs. 55.12 for StreamingVLM) and clips (63.45 vs. 56.22 for DeepSeek-V4.1-Flash). Event-localized credit also raises event recall from 34.5 to 42.1 over supervised fine-tuning. Code and data will be released, and demos are available on an anonymous project page at https://moba-vl.github.io.
Figures & tables
Figure 1: Real-time commentary by MOBA-VL. Given a 1 FPS video stream, the model generates one commentary turn per second using only frames observed so far. Four selected turns from a ten-second sequence are shown. Gray boxes highlight relevant visual evidence and green text marks player references and actions. MOBA-VL describes the elimination at 08:39, while the prior method does not mention it in the displayed turn.
Figure 2: The MOBACast data pipeline. Word-level transcripts are aligned to one-second turns, corrected using match metadata, and filtered for relevance. Removed speech becomes silence on the original timeline.
Figure 3: Event localization and scoring. Given a hypothesis from game telemetry, the NLI locator scores successive commentary prefixes against the event hypothesis. The turn with the largest gain in support is selected as the anchor for credit assignment. The NLI scorer evaluates the full commentary against the same hypothesis to produce the support score used in the reward.
Figure 4: Credit assignment strategies. Top: a commentary timeline where BrokenBlade’s death occurs at 18 s but is described at 21 s. Bottom: token-level advantages assigned by four strategies. Sequence GRPO assigns uniform advantage. Telemetry anchor centers the Gaussian at the event time ( te=18 ), missing the actual description. Event-localized methods use the NLI anchor ( t∗=21 ) and apply either uniform or Gaussian weighting. All variants preserve the mean advantage and bound corrections to 20% of the sequence-level advantage.
Model
Params
League of Legends
Dota 2
Honor of Kings
Overall
CC ↑
Flu ↑
DA ↑
Overall ↑
CC ↑
Flu ↑
DA ↑
Overall ↑
CC ↑
Flu ↑
DA ↑
Overall ↑
Proprietary Models
GPT-5.6-Luna
–
55.0
52.5
50.4
52.6
48.0
56.0
51.5
51.8
51.1
54.7
53.0
52.9
52.47
Claude-Haiku-4.5
–
40.3
50.4
50.5
47.1
36.7
46.0
49.5
44.1
38.3
48.3
52.7
46.5
45.87
Qwen3.7-Plus
–
50.3
47.2
50.4
49.3
44.7
44.0
51.5
46.7
41.9
41.1
51.0
44.7
46.89
Qwen3.8-Flash
–
46.4
43.8
52.6
47.6
37.3
43.3
49.9
43.5
45.8
44.4
50.9
47.0
46.05
Table 1: Main results of full-match evaluation on MOBACast-Bench. For each game, Overall is the arithmetic mean of CC, Flu, and DA, and the final Overall is the arithmetic mean across the three games. Bold indicates the best result, and underline indicates the second best.
Table 3: Pairwise win rates on MOBACast-Bench. Each entry reports the percentage of pairwise comparisons in which Model A is preferred over Model B. The protocol follows StreamingVLM and Proact-VL ( Xu et al., 2026 ; Yan et al., 2026 ) . Bold indicates the best result, and underline indicates the second best.
Method
Recall
GT Hits
Precision
Fluency
SFT base
34.5
100.7
79.5
88.1
+ Sequence RL
38.6
110.4
79.3
88.0
+ Telemetry anchor
38.8
113.2
79.3
88.6
+ Event Localized (Gaussian)
40.8
117.8
79.2
86.9
+ Event Localized (Uniform)
42.1
123.6
79.9
88.6
Table 4: Credit-assignment ablations after 100 updates. Results average three decoding seeds and, for RL variants, three training runs. Precision and Recall are multiplied by 100 and macro-averaged over the six clip subsets. GT Hits counts matched reference events out of 300. Fluency (1–5) is multiplied by 20. Bold/underline indicate best/second-best displayed means.
Method
Video covered
Latency (ms)
RTF
Peak Memory (GiB)
LiveCC
31% (OOM)
20,129 † / 40,281 †
18.35 †
OOM
StreamingVLM
100%
790 / 1,234
0.82
17.3
Proact-VL
100%
459 / 607
0.56
17.0
MOBA-VL
100%
559 / 812
0.60
20.5
w/o cross-turn reuse
100%
6,008 / 10,464
6.08
37.1
w/o periodic reset
76% (OOM)
2,990 † / 5,407 †
2.96 †
OOM
Table 5: Streaming efficiency on a 50-minute match. Latency is the median / P95 turn latency. Memory is peak allocated GPU memory. LiveCC and w/o periodic reset run out of memory after 15.9 and 38.3 minutes. † Computed only over the covered part of the match.
Model
Overall
Buttons
Left stick
Right stick
NitroGen
27.50
22.98
98.67
3.81
Qwen3.5-9B + NitroGen
27.16
22.37
101.09
3.61
MOBA-VL + NitroGen (ours)
21.33
17.39
80.87
3.18
Table 6: Open-loop action prediction on the held-out test set. We calculate the MSE score ( ×10−3 ) for evaluation.
Rubric
NitroGen
Qwen3.5-9B + NitroGen
MOBA-VL + NitroGen (ours)
Equipment purchase ↑
34.00
49.00
49.00
Skill upgrade ↑
40.00
38.00
66.00
Control smoothness ↑
45.00
37.00
40.00
Movement strategy ↑
24.00
28.00
31.00
Attack timing ↑
30.00
35.00
35.00
Teammate support ↑
50.00
49.00
52.00
Table 7: Closed-loop evaluation of the first 130 seconds of gameplay against in-game bots. Scores are reported on a 0–100 scale (higher is better).
Appendix figures & tables12 assets
Supplementary material from the paper’s appendix.
Appendix
Figure 5: SFT training loss. The curve shows the raw values logged every 100 updates, without smoothing.
Setting
SFT
RL
Training windows
412,858
1,349
Validation windows
—
72
Window length
8 s
32 s
Frame rate
1 FPS
1 FPS
Checkpoint step
51,608
100
Windows per update
8
4
Appendix
Table 8: SFT and RL training configurations. Sampling settings refer to RL training rollouts. RL results use the checkpoint at update 100.
Figure 6: Reward curves for Uniform run 1. The left panel shows per-update training rewards and their trailing 10-update mean. The right panel shows validation rewards without smoothing. The panels use different vertical ranges to show their variation. Both use the scalar reward in Eq. equation 1 , including length and repetition penalties.
Setting
Scorer
Locator
Learning rate
2×10−5
5×10−6
Epochs
3
3
Per-GPU batch
8 pairs
Up to 256 prefixes
Gradient accumulation
1
1
GPUs
4
4
Validation interval
500 updates
40 updates
Appendix
Table 9: NLI training settings.
Figure 7: ChatML template for a speaking turn. Frame indices and N , the number of image placeholders, are explanatory annotations. The empty <think> block is supplied by the template before generation and is absent from finalized historical text. Green boxes show newly generated tokens for a turn ending with <|im_end|> ; the final newline is appended by the runtime.
Model
Context policy
Seeds
MOBA-VL
Reset every 32 s; retain four frame–commentary turns
41–43 / 41–43
StreamingVLM
16 s sliding window
41–43 / 41–43
LiveCC
Reset every 75 s; no frame or text carry-over
41–43 / 41–43
Appendix
Table 10: Context policies and reported decoding seeds for the streaming models with explicitly recorded reset or window settings. Seeds are listed as clips / full matches.
Figure 8: Median and P95 whole-turn latency for four configurations, including frame reading and turn completion. The dashed line marks the one-second budget. The horizontal axis is logarithmic. † The variant without periodic reset runs out of memory after 38.3 minutes (76% of the match); its statistics cover only that portion.
Figure 9: Real-time factor and peak allocated GPU memory for the same four configurations. RTF is processing time divided by processed video duration; the dashed line marks RTF 1. † The variant without periodic reset covers only the first 38.3 minutes (76%) of the match.
Figure 10: Mean recorded time per turn for image preprocessing and input construction (Build), vision encoding and LLM prefill (Prefill), token generation (Decode), and output finalization and cache maintenance (Finalize). Position-ID computation rounds to zero milliseconds and is omitted. Component means are rounded independently; total labels use the reported means. † The variant without periodic reset uses only its completed 38.3-minute portion.
Method
Events
Exact ↑
±1 s ↑
MAE ↓
Telemetry timestamp
360
6.4
20.3
–
Scorer as locator
363
47.9
57.3
4.41
Specialized locator
363
56.7
66.4
3.34
Scorer + locator
363
57.9
67.5
3.20
Appendix
Table 11: Anchor agreement with automatic reference annotations on SFT outputs. Exact and ±1 s agreement are percentages; MAE is in seconds. Learned variants use 363 events without detection gating; the telemetry row uses a 360-event diagnostic subset. A dash denotes an unreported metric.
Figure 11: Open-loop MSE at each prediction step within the 16-step action horizon.
Figure 12: Input dependence of VLM latent tokens for five test images. (a–b) Each cell shows the cosine distance between the corresponding token vectors for one image pair, using a shared color scale. (c) Mean distance at each token position across the ten image pairs. Larger distances indicate greater variation across input images.
Real-time video commentary generation provides textual descriptions of ongoing events in videos. It supports accessibility and engagement in domains such as sports, esports, and livestreaming. Commentary generation involves two essential decisions: what to say and when to say it. While recent prompting-based approaches using multimodal large language models (MLLMs) have shown strong performance in content generation, they largely ignore the timing aspect. We investigate whether in-context prompting alone can support real-time commentary generation that is both semantically relevant and well-timed. We propose two prompting-based decoding strategies: 1) a fixed-interval approach, and 2) a novel dynamic interval-based decoding approach that adjusts the next prediction timing based on the estimated duration of the previous utterance. Both methods enable pause-aware generation without any fine-tuning. Experiments on Japanese and English datasets of racing and fighting games show that the dynamic interval-based decoding can generate commentary more closely aligned with human utterance timing and content using prompting alone. We release a multilingual benchmark dataset, trained models, and implementations to support future research on real-time video commentary generation.
Anum Afzal, Yuki Saito, Hiroya Takamura +5
School of CIT, Technical University of Munich · National Institute of Advanced Industrial Science and Technology (AIST) · The University of Tokyo +3
Given the rapidly growing capabilities of vision-language models (VLMs), extending them to interactive decision-making tasks such as video games has emerged as a promising frontier. However, existing approaches either rely on large-scale supervised fine-tuning (SFT) on human trajectories or apply reinforcement learning (RL) only in relatively short-horizon settings (typically around 20--30 turns). In this work, we study RL-based training of VLMs for long-horizon decision-making in Super Mario Land, a visually grounded environment requiring 100+ turns of interaction with coordinated perception, reasoning, and action. We begin with a systematic investigation of key algorithmic components and propose an adapted variant of PPO with a lightweight turn-level critic, which substantially improves training stability and sample efficiency over critic-free methods such as GRPO and Reinforce++. We further show that pretrained VLMs provide strong action priors, significantly improving sample efficiency during RL training and reducing the need for manual design choices such as action engineering, compared to classical deep RL trained from scratch. Building on these insights, we introduce Odysseus, an open training framework for VLM agents, achieving substantial gains across multiple levels of the game and at least 3 times average game progresses than frontier models. Moreover, the trained models exhibit consistent improvements under both in-game and cross-game generalization settings, while maintaining general-domain capabilities. Overall, our results identify key ingredients for making RL stable and effective in long-horizon, multi-modal settings, and provide practical guidance for developing VLMs as embodied agents.
Chengshuai Shi, Wenzhe Li, Xinran Liang +10
Princeton Language and Intelligence, Princeton University · Fudan University · Tsinghua University
We present Wan-Streamer v0.3, which reframes our native-streaming interaction model under a single organizing view: a video is a world plus an event stream. The world is the persistent context in which a video unfolds, including the environment, scene, subjects, ambient acoustic conditions, voice characteristics, and other relatively stable conditions. The event stream is everything that changes over time within that world, including scene or environmental changes, subject behavior, speech, and other sounds. This yields a general-purpose pretraining task over large amounts of real video: given a world and incoming input, predict how the world moves, changes, and responds in real time. The resulting competence can be specialized to a broad family of real-time downstream tasks. We instantiate it on real-time full-duplex audio-visual interaction, where the event stream is the agent's speech together with free-form behavior. Functionally, the model's multimodal understanding process is vision-language-action-like: it maps multimodal user input to language-form speech and behavior actions. Wan-Streamer v0.3 preserves the v0.2 operating point: 640x368 video at 25 FPS, a 160 ms streaming unit, approximately 200 ms model-side response latency, and approximately 550 ms total interaction latency under a 350 ms bidirectional network budget.