cs.CVSep 30, 2026

VR-JEPA: Learning Contrastive-State Latent Guidance for Generation-based Video Reasoning

Authors: Zehua Ma, Kun Xiang, Yunshuang Nie, Quanlin Chen, Haoyuan Li, Xiuwei Chen, Jiang Ji, Haijun Wu, +4 more

Organizations: Shenzhen Campus of Sun Yat-sen University · Tsinghua Shenzhen International Graduate School · Shenzhen Loop Area Institute · Tencent · Mohamed bin Zayed University of Artificial Intelligence · UiT The Arctic University of Norway

Abstract

Reasoning through video generation offers a promising path toward visual intelligence by modeling latent visual states and their dynamics. However, current video generation models often lack explicit guidance on how these states should evolve, leaving generated trajectories prone to physical and structural inconsistencies that undermine reasoning reliability. While the Video Joint-Embedding Predictive Architecture (V-JEPA) provides rich spatiotemporal priors learned through latent prediction, these general priors do not naturally adapt to the logical reasoning capabilities required for complex visual tasks. To bridge this gap, we propose VR-JEPA, a framework that aligns the V-JEPA predictor with task-specific reasoning logic through localized contrastive-state learning and uses its predicted latent trajectories to guide video generation for visual reasoning. Specifically, (i) we pair successful trajectories with generated alternatives under the same input conditions and use discrepancies in their V-JEPA representations to identify informative states and tokens for localized contrastive supervision. (ii) We further equip the V-JEPA predictor with skill-specific experts trained on anchor-task data, allowing the model to adaptively specialize its shared spatiotemporal priors across diverse cognitive domains. Together with skill-specific experts, this contrastive supervision enables VR-JEPA to predict latent trajectories that provide task-specific logical guidance for video generation. Comprehensive experiments on the large-scale VBVR-Pro-Bench dataset demonstrate that VR-JEPA achieves an 11.33%11.33\% relative improvement over the cutting-edge generation-based reasoning baseline, significantly mitigating physical artifacts and enhancing logical consistency.

Figures & tables

Appendix figures & tables10 assets

Supplementary material from the paper’s appendix.

Appendix

Explore similar work

CardsList
  1. From Evaluation to Enhancement: Benchmarking and Improving Think-with-Video Reasoning for Video Generative Models

    Sep 12, 2026Meng Luo, Yicheng Liu, Jiahao Wang +5Generative Video ModelsVideo Understanding

  2. OpenCoF: Learning to Reason Through Video Generation

    Jul 9, 2026Xinyan Chen, Ziyu Guo, Renrui Zhang +2Video GenerationVideo Understanding

  3. CollabVR: Collaborative Video Reasoning with Vision-Language and Video Generation Models

    May 9, 2026Joowon Kim, Seungho Shin, Joonhyung Park +1Video UnderstandingVisual Reasoning