Token-Level Video Reinforcement Learning
Organizations: Northeastern University
Abstract
Reinforcement learning (RL) for video generation usually assigns one scalar reward to an entire sampled video. Yet a video is not uniformly flawed: some visual tokens may already satisfy the prompt, whereas others require correction. A scalar reward cannot localize errors, causing optimization to perturb satisfactory tokens while under-targeting the tokens that actually need to change. We introduce Token-Level Video Reinforcement Learning, TVRL, a framework that derives token-level credit from the reward being optimized. Our key insight is that the answer likelihood of a frozen vision-language model provides both signals: its outputs contribute to the video-level reward, while magnitudes of its video-input gradients reveal which generated video tokens most affect that score. We instantiate TVRL in Group Relative Policy Optimization by averaging prompt-derived question rewards into one group-relative advantage and using detached, question-conditioned token-credit maps to reweight dense denoising-transition log-probabilities inside the clipped policy ratio. On VBench-2.0, TVRL achieves an Overall score of 57.69, outperforming the base model by 3.60 points. TVRL also improves matched GRPO baselines across three SDE samplers (SAGE, Flow, and Dance) by 2.68--3.15 points and across four reward models (VideoAlign, VideoScore2, UnifiedReward2, and Qwen3.5-9B) by 1.33--3.15 points.
Figures & tables
| Pair | Win | Loss | Tie |
|---|---|---|---|
| vs. GRPO (SAGE) | 36.3 | 23.3 | 40.4 |
| vs. Base | 31.3 | 19.2 | 49.5 |
Appendix figures & tables11 assets
Supplementary material from the paper’s appendix.
Appendix
| Reward | Method | Overall | Creat. | Comm. | Ctrl. | Human | Phys. | |
|---|---|---|---|---|---|---|---|---|
| Base model | – | 54.09 0.11 | 41.40 0.38 | 62.75 0.24 | 30.26 0.46 | 88.94 0.13 | 47.11 0.34 | 3 |
| VideoAlign | GRPO | 54.18 0.23 | 41.44 0.41 | 61.14 0.26 | 30.77 0.44 | 90.06 0.12 | 47.49 0.35 | 3 |
| TVRL | 55.54 0.42 | 45.11 0.36 | 61.16 0.27 | 32.09 0.43 | 90.21 0.14 | 49.15 0.32 | 3 | |
| VideoScore2 | GRPO | 54.66 0.16 | 42.23 0.39 | 64.89 0.22 | 30.29 0.48 | 91.52 0.11 | 44.35 0.37 | 3 |
| TVRL | 55.99 0.33 | 42.08 0.40 | 64.60 0.21 | 31.33 0.45 | 90.79 0.13 | 49.13 0.31 | 3 | |
| UnifiedReward2 | GRPO | 54.82 0.47 | 42.90 0.37 | 62.14 0.25 | 30.25 0.49 | 88.90 0.15 | 50.90 0.30 | 3 |
| Setting | Overall | |
|---|---|---|
| Sampler | SAGE, GRPO | 54.54 0.38 |
| SAGE, TVRL | 57.69 0.13 | |
| Flow, GRPO | 53.79 0.31 | |
| Flow, TVRL | 56.71 0.27 | |
| Dance, GRPO | 50.84 0.42 | |
| Dance, TVRL | 53.52 0.34 |
| Variant | Reward | Credit | Overall | Creat. | Comm. | Ctrl. | Human | Phys. |
|---|---|---|---|---|---|---|---|---|
| Free generation | Free-gen | Uniform | 54.28 | 37.32 | 67.75 | 29.75 | 87.42 | 49.14 |
| Log-probability reward | Logprob | Uniform | 54.54 | 41.68 | 64.88 | 31.57 | 88.85 | 45.74 |
| Token credit | Logprob | Per-question gradient | 54.42 | 39.98 | 66.31 | 30.38 | 88.41 | 47.02 |
| Control | Setting | Overall | Creat. | Comm. | Ctrl. | Human | Phys. |
|---|---|---|---|---|---|---|---|
| Rating threshold | 5 | 55.41 | 36.98 | 69.20 | 33.03 | 88.71 | 49.14 |
| Reward frames | 5 | 53.49 | 40.81 | 65.47 | 30.42 | 87.08 | 43.64 |
| 30 | 55.71 | 44.60 | 61.37 | 30.41 | 90.02 | 52.14 | |
| Questions | 1 | 51.34 | 33.52 | 64.34 | 27.91 | 87.62 | 43.30 |
| 3 | 55.69 | 40.24 | 65.43 | 31.58 | 90.20 | 51.00 |
| Ablation | Overall | Creat. | Comm. | Ctrl. | Human | Phys. |
|---|---|---|---|---|---|---|
| General semantic credit | 54.96 | 44.00 | 66.31 | 27.41 | 88.46 | 48.64 |
| Averaged map, single ratio | 54.96 | 42.40 | 64.88 | 34.71 | 86.94 | 45.88 |
| Per-frame ratio | 52.95 | 38.77 | 66.03 | 29.04 | 84.66 | 46.26 |
| Method | Mean (s) | Median (s) | Relative time |
|---|---|---|---|
| GRPO | 666.05 | 661.72 | |
| TVRL | 786.04 | 781.53 |
| Panels | Cohort | Pairs | (pp), 95% CI | Queried mass |
|---|---|---|---|---|
| Crop-to-fill | Prequalified | 20 | 52.70% | |
| Crop-to-fill | All pairs | 24 | 53.20% | |
| Letterbox | Prequalified | 23 | 51.90% | |
| Letterbox | All pairs | 24 | 51.87% |
| Component | Setting |
|---|---|
| Generator | Text-to-video diffusion transformer, pretrained 480p_t2v checkpoint |
| Post-training objective | GRPO with TVRL token credit |
| Main SDE type | sage_grpo |
| Reward critic | Frozen Qwen3.5-9B VLM |
| Reward model family | qwen3_5 |
| Reward score type | token_credit |
| SDE | Method | Overall | |
|---|---|---|---|
| SAGE | GRPO | 54.18 | – |
| TVRL | 55.54 | +1.36 | |
| Flow | GRPO | 55.07 | – |
| TVRL | 55.82 | +0.75 | |
| Dance | GRPO | 48.02 | – |
| TVRL | 50.30 | +2.28 |