Reinforcement learning (RL) for video generation usually assigns one scalar reward to an entire sampled video. Yet a video is not uniformly flawed: some visual tokens may already satisfy the prompt, whereas others require correction. A scalar reward cannot localize errors, causing optimization to perturb satisfactory tokens while under-targeting the tokens that actually need to change. We introduce Token-Level Video Reinforcement Learning, TVRL, a framework that derives token-level credit from the reward being optimized. Our key insight is that the answer likelihood of a frozen vision-language model provides both signals: its outputs contribute to the video-level reward, while magnitudes of its video-input gradients reveal which generated video tokens most affect that score. We instantiate TVRL in Group Relative Policy Optimization by averaging prompt-derived question rewards into one group-relative advantage and using detached, question-conditioned token-credit maps to reweight dense denoising-transition log-probabilities inside the clipped policy ratio. On VBench-2.0, TVRL achieves an Overall score of 57.69, outperforming the base model by 3.60 points. TVRL also improves matched GRPO baselines across three SDE samplers (SAGE, Flow, and Dance) by 2.68--3.15 points and across four reward models (VideoAlign, VideoScore2, UnifiedReward2, and Qwen3.5-9B) by 1.33--3.15 points.
Figures & tables
Figure 1: Text-to-video RL post-training with TVRL. Compared with the base model (top rows), TVRL (bottom rows) improves subject presence, object fidelity, body plausibility, and camera-motion control; red marks the prompt phrase to check.
Figure 2: Video-level reward, token-level credit. Video GRPO ( Zheng et al., 2026 ) broadcasts one scalar advantage to every video token. TVRL provides token-level credit to acknowledge or penalize relevant tokens.
Figure 3: Overview of TVRL. Teacher-forced likelihoods for K prompt-derived checks are averaged within each of M rollouts and normalized across the rollout group into one advantage per video. Gradients of the same likelihoods yield detached, question-conditioned weights that route dense denoising log-probabilities inside the clipped GRPO update.
Figure 4: Qualitative results on HunyuanVideo-1.5. Comparison of prompt- and seed-matched videos from the base model ( Wu et al., 2025a ) , Dance-GRPO ( Xue et al., 2025 ) , Flow-GRPO ( Liu et al., 2025a ) , SAGE-GRPO ( Zheng et al., 2026 ) , and TVRL with the Qwen3.5-9B critic on SAGE-GRPO validation prompts (full text in Section A.2 ); row labels name the prompt requirement to check. Only TVRL satisfies all three, whereas the baselines substitute a flat zither, put a fist or an object in the raised hand, or lose the low-angle framing and the dress.
Figure 5: VBench-2.0 Overall of GRPO and TVRL across SDE samplers.
Pair
Win
Loss
Tie
vs. GRPO (SAGE)
36.3
23.3
40.4
vs. Base
31.3
19.2
49.5
Table 2: Blind pairwise human preferences for TVRL against the matched GRPO baseline under SAGE and the base model (%).
Figure 6: Effect of the frozen critic. Prompt- and seed-matched videos from the base model ( Wu et al., 2025a ) and TVRL with VideoAlign, VideoScore2, Qwen3.5-4B, or Qwen3.5-9B as the critic on VideoGen-Eval prompts ( Yang et al., 2025 ) ( Section A.2 ); row labels name the aspect to check. Only the Qwen3.5-9B critic keeps all three, whereas weaker critics add a third person, show the couple from behind, or lose the road.
Figure 7: Finer credit learns faster. Training reward with Qwen3.5-9B (left) and VideoScore2 (right), averaged over three runs: w/o credit < frame <7×7<3×3 later in training.
Table 9
Figure 8: Visualization of residuals Δ(1×1,3×3) and Δ(3×3,7×7) .
Appendix figures & tables11 assets
Supplementary material from the paper’s appendix.
Appendix
Reward
Method
Overall
Creat.
Comm.
Ctrl.
Human
Phys.
n
Base model
–
54.09 ± 0.11
41.40 ± 0.38
62.75 ± 0.24
30.26 ± 0.46
88.94 ± 0.13
47.11 ± 0.34
3
VideoAlign
GRPO
54.18 ± 0.23
41.44 ± 0.41
61.14 ± 0.26
30.77 ± 0.44
90.06 ± 0.12
47.49 ± 0.35
3
TVRL
55.54 ± 0.42
45.11 ± 0.36
61.16 ± 0.27
32.09 ± 0.43
90.21 ± 0.14
49.15 ± 0.32
3
VideoScore2
GRPO
54.66 ± 0.16
42.23 ± 0.39
64.89 ± 0.22
30.29 ± 0.48
91.52 ± 0.11
44.35 ± 0.37
3
TVRL
55.99 ± 0.33
42.08 ± 0.40
64.60 ± 0.21
31.33 ± 0.45
90.79 ± 0.13
49.13 ± 0.31
3
UnifiedReward2
GRPO
54.82 ± 0.47
42.90 ± 0.37
62.14 ± 0.25
30.25 ± 0.49
88.90 ± 0.15
50.90 ± 0.30
3
Appendix
Table A1: Table 1 with standard deviations over seeds (SAGE sampler, step 100).
Setting
Overall
Sampler
SAGE, GRPO
54.54 ± 0.38
SAGE, TVRL
57.69 ± 0.13
Flow, GRPO
53.79 ± 0.31
Flow, TVRL
56.71 ± 0.27
Dance, GRPO
50.84 ± 0.42
Dance, TVRL
53.52 ± 0.34
Appendix
Table A2: Overall mean ± standard deviation for the sampler comparison ( Figure 5 ), the routing ablation ( Table 4 ), and the critic ablation ( Table 4 ), all with the Qwen3.5-9B reward. Dance GRPO averages two runs; all other rows average three.
Variant
Reward
Credit
Overall
Creat.
Comm.
Ctrl.
Human
Phys.
Free generation
Free-gen
Uniform
54.28
37.32
67.75
29.75
87.42
49.14
Log-probability reward
Logprob
Uniform
54.54
41.68
64.88
31.57
88.85
45.74
Token credit
Logprob
Per-question gradient
54.42
39.98
66.31
30.38
88.41
47.02
Appendix
Table A3: Reward-objective ablations preceding the final shared-advantage setting. All rows use SAGE-GRPO with Qwen3.5-9B and the same decomposed binary-question prompts, while varying VLM scoring mode and temporal credit assignment.
Control
Setting
Overall
Creat.
Comm.
Ctrl.
Human
Phys.
Rating threshold
5
55.41
36.98
69.20
33.03
88.71
49.14
Reward frames
5
53.49
40.81
65.47
30.42
87.08
43.64
30
55.71
44.60
61.37
30.41
90.02
52.14
Questions
1
51.34
33.52
64.34
27.91
87.62
43.30
3
55.69
40.24
65.43
31.58
90.20
51.00
Appendix
Table A4: Rating-threshold, reward-frame-count, and question-count ablations for SAGE-GRPO with Qwen3.5-9B token credit, evaluated after 100 optimization steps. All score columns report VBench-2.0 results.
Ablation
Overall
Creat.
Comm.
Ctrl.
Human
Phys.
General semantic credit
54.96
44.00
66.31
27.41
88.46
48.64
Averaged map, single ratio
54.96
42.40
64.88
34.71
86.94
45.88
Per-frame ratio
52.95
38.77
66.03
29.04
84.66
46.26
Appendix
Table A5: Ablations on how VLM-gradient credit is constructed and applied, for SAGE-GRPO with Qwen3.5-9B evaluated after 100 optimization steps. All rows use frame-level credit; per-question frame credit reaches 56.47. General semantic credit uses one generic check; averaged map uses one ratio with wˉ=K1∑jwj , the first-order equivalent of per-question ratios; per-frame ratio clips frame-wise ratios before aggregation.
Method
Mean (s)
Median (s)
Relative time
GRPO
666.05
661.72
1.000×
TVRL
786.04
781.53
1.180×
Appendix
Table A6: Observed training overhead in historical SAGE/VideoAlign runs. Times are seconds per update over updates 11–100; relative time is the ratio of means. Configuration differences are described in Section A.5 .
Panels
Cohort
Pairs
D (pp), 95% CI
Queried mass
Crop-to-fill
Prequalified
20
+5.39[2.83,7.77]
52.70%
Crop-to-fill
All pairs
24
+6.41[4.46,8.43]
53.20%
Letterbox
Prequalified
23
+3.79[1.66,5.99]
51.90%
Letterbox
All pairs
24
+3.74[1.63,5.90]
51.87%
Appendix
Table A7: Question–position crossover diagnostic with Qwen3.5-9B at the training critic resolution. D measures question-induced panel-credit change in percentage points (pp); the final column is mean credit mass in the queried panel, which equals 50% under uniform routing. Every cohort contains 12 prompt-pair groups. Crop-to-fill is the primary panel construction; letterbox is the earlier construction on the same pairs and questions.
Table B2: VBench-2.0 Overall with VideoAlign as the reward model across SDE samplers; Δ is the gain of TVRL over GRPO.
Figure B1: Supplementary optimization traces including the unsmoothed 1×1 credit map for Qwen3.5-9B (left) and VideoScore2 (right). The cell-level variant does not continue the improvement obtained by progressively localizing spatially aggregated credit and is particularly unstable with VideoScore2.
Figure B2: Additional spatial residual diagnostics. The positive 3×3−7×7 residual for two further question-conditioned maps, shown over five sampled frames (queried evidence in red). At high-credit frames, the strongest residuals lie near the pen and gold-rimmed barrel and near the white headband. These per-example normalized sensitivity maps are not causal masks.
Video diffusion models have made rapid progress in perceptual realism and temporal coherence, but they remain primarily optimized for plausible generation rather than verifiable reasoning. This limitation is especially pronounced in tasks where generated videos must satisfy explicit spatial, temporal, or logical constraints. Inspired by the role of reinforcement learning with verifiable rewards (RLVR) in reasoning-oriented language models, we introduce VideoRLVR, a practical recipe for optimizing video diffusion models with rule-based feedback. VideoRLVR formulates video reasoning as the generation of verifiable visual trajectories and consists of an SDE-GRPO optimization backbone, dense decomposed rewards, and an Early-Step Focus strategy for efficient training. The Early-Step Focus strategy restricts policy optimization to the early denoising phase, reducing training latency by about 40% while preserving performance. We evaluate VideoRLVR on Maze, FlowFree, and Sokoban, three procedurally generated domains with objective success criteria. Across these tasks, VideoRLVR consistently improves over supervised fine-tuning baselines, with dense decomposed rewards proving especially important in low-success-rate settings. Our RL-optimized model also outperforms the evaluated proprietary and open-source video generation models on these verifiable reasoning benchmarks and out-of-domain benchmarks. These results suggest that verifiable RL can move video models beyond perceptual imitation toward more reliable rule-consistent visual reasoning.
Tinghui Zhu, Sheng Zhang, James Y. Huang +5
University of California, Davis · Microsoft Research · University of Southern California +1
Reinforcement learning from verifiable rewards (RLVR) has demonstrated remarkable effectiveness in improving the reasoning capabilities of large language models. As models evolve into natively multimodal architectures, extending RLVR to video understanding becomes increasingly important yet remains largely unexplored, due to the diversity of video task types, the computational overhead of repeatedly decoding and preprocessing high-dimensional visual inputs, and the difficulty of reproducible evaluation across numerous sensitive hyperparameters. Existing open-source RL training frameworks provide solid infrastructure for text and image scenarios but lack systematic optimizations tailored for video modality. In this work, we present \textbf{EasyVideoR1}, a complete and efficient reinforcement learning framework specifically designed for training large vision-language models on video understanding tasks. EasyVideoR1 makes the following contributions: (1) a full video RL training pipeline with offline preprocessing and tensor caching that eliminates redundant video decoding and yields a 1.47 × throughput improvement; (2) a comprehensive, task-aware reward system covering 11 distinct video and image problem types with unified routing and modular extension; (3) a mixed offline-online data training paradigm that combines curated high-quality trajectories with on-policy exploration, benefiting the learning of more challenging tasks; (4) joint image-video training with independently configurable pixel budgets, allowing the two modalities to mutually reinforce each other; and (5) an asynchronous multi-benchmark evaluation framework covering 22 mainstream video understanding benchmarks, with reproduced accuracy closely aligned with officially reported scores.
Chuanyu Qin, Chenxu Yang, Qingyi Si +6
Institute of Information Engineering, Chinese Academy of Sciences, Beijing, China · School of Cyber Security, University of Chinese Academy of Sciences, Beijing, China
Reward models for text-to-video (T2V) generation guide post-training but often fail at fine-grained semantic alignment. We trace this to two structural weaknesses in existing reasoning-based reward models: they do not systematically verify every condition described in the prompt, and the visual evidence supporting each judgment remains implicit in their free-form reasoning. We propose SG-PVR, a video reward model that addresses these limitations through plan-and-verify reasoning grounded in spatio-temporal scene graphs. The verification plan decomposes the prompt into atomic claims, making the set of requirements to be checked explicit. The spatio-temporal scene graph, encoding entities, attributes, and temporally grounded relations, is extracted from the video and maintained as a persistent structured visual reference throughout reasoning. Each claim is verified against both the video and the scene graph, anchoring judgments in explicit visual evidence. SG-PVR achieves strong performance on semantic alignment, including fine-grained temporal semantics. As a test-time reranker, it further enhances compositional alignment in T2V generation.
Hyomin Kim, Junghye Kim, Joanie Hayoun Chung +4
Department of Artificial Intelligence, Korea University · Department of Statistics, Korea University