cs.CVOct 1, 2026

Token-Level Video Reinforcement Learning

Authors: Yifan Wang, Gordon Guocheng Qian, Yanyu Li, Anil Kag, Yun Fu

Organizations: Northeastern University

Abstract

Reinforcement learning (RL) for video generation usually assigns one scalar reward to an entire sampled video. Yet a video is not uniformly flawed: some visual tokens may already satisfy the prompt, whereas others require correction. A scalar reward cannot localize errors, causing optimization to perturb satisfactory tokens while under-targeting the tokens that actually need to change. We introduce Token-Level Video Reinforcement Learning, TVRL, a framework that derives token-level credit from the reward being optimized. Our key insight is that the answer likelihood of a frozen vision-language model provides both signals: its outputs contribute to the video-level reward, while magnitudes of its video-input gradients reveal which generated video tokens most affect that score. We instantiate TVRL in Group Relative Policy Optimization by averaging prompt-derived question rewards into one group-relative advantage and using detached, question-conditioned token-credit maps to reweight dense denoising-transition log-probabilities inside the clipped policy ratio. On VBench-2.0, TVRL achieves an Overall score of 57.69, outperforming the base model by 3.60 points. TVRL also improves matched GRPO baselines across three SDE samplers (SAGE, Flow, and Dance) by 2.68--3.15 points and across four reward models (VideoAlign, VideoScore2, UnifiedReward2, and Qwen3.5-9B) by 1.33--3.15 points.

Figures & tables

Appendix figures & tables11 assets

Supplementary material from the paper’s appendix.

Appendix

Explore similar work

CardsList
  1. Video Models Can Reason with Verifiable Rewards

    May 14, 2026Tinghui Zhu, Sheng Zhang, James Y. Huang +5Video Diffusion ModelsReinforcement Learning With Verifiable Reward

  2. EasyVideoR1: Easier RL for Video Understanding

    Apr 18, 2026Chuanyu Qin, Chenxu Yang, Qingyi Si +6Video UnderstandingOffline Reinforcement Learning

  3. Plan-and-Verify Video Reward Reasoning with Spatio-Temporal Scene Graph Grounding

    Jun 10, 2026Hyomin Kim, Junghye Kim, Joanie Hayoun Chung +4Text-To-Video Generation ModelSpatio-Temporal Scene Graphs