Organizations: Nankai University · Carnegie Mellon University · Xiaohongshu Inc. · Columbia University · Santa Clara University · New York University · University of California, Santa Cruz
Multimodal Large Language Models (MLLMs) have demonstrated remarkable potential in video understanding, yet their reliance on retrospective summarization and text-centric priors often limits their ability to bridge unobserved causal transitions when applied to Video Event Prediction (VEP). To address this, we propose VepAgent, an agentic framework that integrates causal-transition reasoning with tool-augmented reinforcement learning (RL) for robust VEP. Unlike prior methods that passively project future trajectories from historical dependencies, our approach explicitly models the logical progression from terminal observed states to future events. Specifically, we first construct futurebench-4K, a high-quality chain-of-thought dataset for supervised fine-tuning (SFT) that effectively bridges the causal-logic gap by structuring the deduction of unobserved intermediate states. Subsequently, we develop a diagnostic tool library integrating state tracking, frame retrieval, and region magnification, enabling the agent to dynamically augment reasoning with external tools to recover missing spatio-temporal evidence and resolve visual ambiguities during inference. Moreover, we propose a composite reward mechanism that jointly optimizes prediction accuracy, causal coherence, and reliable prior, compelling the agent to rely on genuine visual grounding rather than superficial textual similarities. Extensive evaluations on FutureBench and NEPBench datasets demonstrate that our method achieves state-of-the-art performance, significantly outperforming larger MLLMs and validating the empirical effectiveness of our agentic, future-oriented reasoning paradigm.
Figures & tables
Figure 1: Key insight of VepAgent. (a) Vanilla VLM reasoning weakly models the future transition and predicts a coarse continuation of the observed action. (b) Option-driven reasoning is distracted by textual priors and prematurely predicts the cooking stage. (c) Causal-gap reasoning captures the ongoing buttering process, but misses the fine-grained visual state and selects the wrong continuation. (d) With tool-augmented visual evidence, VepAgent identifies the remaining unbuttered slice, correctly reasons over the unfinished action, and predicts the next event.
Figure 2: Pipeline of VepAgent. (i) Construct tool-augmented causal-transition CoT data by explicitly modeling the progression from terminal observed states to future events. (ii) Supervised fine-tuning initializes tool-use and causal-gap reasoning behaviors. (iii) GRPO optimizes the policy with accuracy, causal-gap, and anti-prior rewards. (iv) The trained agent dynamically integrates diagnostic tools for grounded future-event prediction.
Figure 3: Diagnostic tool library of VepAgent. FRE retrieves dense visual evidence from selected temporal windows, while RFM crops and magnifies task-relevant spatial regions. STT summarizes observed state trajectories and the terminal state, whereas VDE extracts fine-grained visible facts as textual evidence. All tools operate exclusively on the observed portion of the video.
Method
Frames
1-Hop
2-Hop
3-Hop
Interp.
AVG
GLM-4.1V-9B
32
29.90
41.90
52.20
47.30
44.38
LLaVA-NeXT-Video
32
48.80
49.30
40.00
44.40
45.24
Kimi-VL-A3B
32
44.30
42.80
51.30
51.90
48.87
MiMo-VL-7B
32
59.00
59.60
50.50
43.80
50.45
InternVL3-8B
32
54.30
58.00
63.20
54.40
56.72
Qwen2.5-VL-32B
32
66.50
62.70
63.20
55.20
59.94
Table 1: Event Prediction Acc. (%) on FutureBench. best and second best results are highlighted.
Method
Overall
Charades
VidSitu
YouCook2
GLM-4.1V-9B
43.73
45.98
28.18
64.48
InternVL3-8B
43.60
37.01
40.55
52.26
Kimi-VL-A3B
41.20
28.66
53.06
32.44
LLaVA-Video-7B
29.60
25.67
24.59
39.32
MiMo-VL-7B
44.53
40.47
40.83
52.46
Qwen2.5-VL-72B
47.50
35.59
41.05
64.48
Table 2: Event Prediction Acc. (%) on NEPBench. best and second best results are highlighted.
Figure 4: Case study between Qwen3-VL-4B and VepAgent-4B. The baseline predicts an already observed meat-addition event. StateTransitionTracker reconstructs the observed progression from meat addition to mixing and active cooking. Grounded in evidence, VepAgent identifies the unfinished cooking process, reasons over subsequent state evolution, and predicts the correct future event.
Method
1-Hop ↑
3-Hop ↑
AVG ↑
Tools
1-Hop ↑
3-Hop ↑
AVG ↑
(a) Qwen3-VL-4B
54.91
59.20
59.09
(e) w/o. STT
82.08
71.64
75.66
(b) 4B + SFT only
77.46
75.62
71.59
(f) w/o. FRE
82.08
75.62
76.99
(c) 4B + SFT + RL
79.77
71.14
75.09
(g) w/o. RFM
82.66
76.12
78.31
(d) 4B + SFT + RL + Tools ( G=4 )
83.82
76.12
78.79
(h) w/o. VDE
83.24
75.12
77.46
Reward Term
1-Hop ↑
3-Hop ↑
AVG ↑
Num of Gen.
1-Hop ↑
3-Hop ↑
AVG ↑
(i) R = Racc
78.61
76.12
74.62
(m) Num = 2
84.39
77.61
78.60
Table 3: Ablation studies (%) on FutureBench. Component-wise ablations are conducted with G=4 . best and second best results are highlighted.
Appendix figures & tables5 assets
Supplementary material from the paper’s appendix.
Appendix
Strategy
1-Hop
3-Hop
AVG
Avg Calls
(a) No Tools
79.77
71.14
75.09
0.00
(b) All Tools (Static)
76.30
70.65
74.91
4.00
(c) Random Tools
82.08
77.61
76.99
1.85
(d) Dynamic (Ours)
83.82
76.12
78.79
1.13
Appendix
Table 4: Ablation on tool selection strategies.
Tool Configuration
1-Hop
3-Hop
AVG
(a) Temporal Only (STT + FRE)
82.08
75.62
77.84
(b) Spatial Only (RFM + VDE)
80.92
71.64
75.38
(c) All (Ours)
83.82
76.12
78.79
Appendix
Table 5: Ablation on tool synergy between temporal and spatial dimensions.
Figure 5: Additional case study with FocusedEvidenceRetriever. The baseline repeatedly predicts the earlier fork-pressing action. FRE retrieves dense late-window visual evidence showing that the fork interaction has ended while the flatbread continues browning. VepAgent uses this temporal evidence to reason toward the subsequent cooking transition.
Figure 6: Additional case study with VisualDetailExtractor. The baseline repeats the already observed onion-addition event. VDE verifies fine-grained visual states, including partially translucent onions being stirred in the pot. VepAgent therefore recognizes that onion addition has completed and reasons toward the next ingredient stage.
Tool
Input
Output
Failure Handling
State Transition Tracker (STT)
observed video
textual event chain & final state
fallback to global video caption
Focused Evidence Retriever (FRE)
video & temporal window
dense frame sequence (contact sheet)
ignored if temporal window is invalid
Region Focus Magnifier (RFM)
video & predicted spatial boxes
magnified image patch
ignored if boxes are invalid
Visual Detail Extractor (VDE)
final frames & scene query
textual fine-grained facts
fallback to original query
Appendix
Table 6: Implementation details of the diagnostic tool library.