cs.CVOct 5, 2026

VepAgent: Bridging Causal-Transition via Tool-Augmented Reinforcement Learning for Video Event Prediction

Authors: Qiutong Chen, Yuchan Guo, Zhenlong Yuan, Haobo Yang, Fangfang Lin, Xinyi Long, Yin Wang, Zijian Song, +5 more

Organizations: Nankai University · Carnegie Mellon University · Xiaohongshu Inc. · Columbia University · Santa Clara University · New York University · University of California, Santa Cruz

Abstract

Multimodal Large Language Models (MLLMs) have demonstrated remarkable potential in video understanding, yet their reliance on retrospective summarization and text-centric priors often limits their ability to bridge unobserved causal transitions when applied to Video Event Prediction (VEP). To address this, we propose VepAgent, an agentic framework that integrates causal-transition reasoning with tool-augmented reinforcement learning (RL) for robust VEP. Unlike prior methods that passively project future trajectories from historical dependencies, our approach explicitly models the logical progression from terminal observed states to future events. Specifically, we first construct futurebench-4K, a high-quality chain-of-thought dataset for supervised fine-tuning (SFT) that effectively bridges the causal-logic gap by structuring the deduction of unobserved intermediate states. Subsequently, we develop a diagnostic tool library integrating state tracking, frame retrieval, and region magnification, enabling the agent to dynamically augment reasoning with external tools to recover missing spatio-temporal evidence and resolve visual ambiguities during inference. Moreover, we propose a composite reward mechanism that jointly optimizes prediction accuracy, causal coherence, and reliable prior, compelling the agent to rely on genuine visual grounding rather than superficial textual similarities. Extensive evaluations on FutureBench and NEPBench datasets demonstrate that our method achieves state-of-the-art performance, significantly outperforming larger MLLMs and validating the empirical effectiveness of our agentic, future-oriented reasoning paradigm.

Figures & tables

Appendix figures & tables5 assets

Supplementary material from the paper’s appendix.

Appendix

Explore similar work

CardsList
  1. Imagine Before You Predict: Interleaved Latent Visual Reasoning for Video Event Prediction

    Jun 4, 2026Tianxiang Jiang, Linquan Wu, Sheng Xia +5Future Video PredictionLatent Visual Reasoning

  2. TEMPURA: Temporal Event Masked Prediction and Understanding for Reasoning in Action

    May 2, 2025Jen-Hao Cheng, Yi-Hao Peng, Huapeng Zhou +11Video Temporal GroundingTemporal Grounding

  3. Towards Effective Long-Video Event Prediction via Multi-Level Event Semantics Mining

    May 29, 2026Bo Peng, YuanJie Lyu, PengGang Qin +1Future Video PredictionVideo-Language Models