Organizations: Nankai University · Carnegie Mellon University · Xiaohongshu Inc. · Columbia University · Santa Clara University · New York University · University of California, Santa Cruz
Multimodal Large Language Models (MLLMs) have demonstrated remarkable potential in video understanding, yet their reliance on retrospective summarization and text-centric priors often limits their ability to bridge unobserved causal transitions when applied to Video Event Prediction (VEP). To address this, we propose VepAgent, an agentic framework that integrates causal-transition reasoning with tool-augmented reinforcement learning (RL) for robust VEP. Unlike prior methods that passively project future trajectories from historical dependencies, our approach explicitly models the logical progression from terminal observed states to future events. Specifically, we first construct futurebench-4K, a high-quality chain-of-thought dataset for supervised fine-tuning (SFT) that effectively bridges the causal-logic gap by structuring the deduction of unobserved intermediate states. Subsequently, we develop a diagnostic tool library integrating state tracking, frame retrieval, and region magnification, enabling the agent to dynamically augment reasoning with external tools to recover missing spatio-temporal evidence and resolve visual ambiguities during inference. Moreover, we propose a composite reward mechanism that jointly optimizes prediction accuracy, causal coherence, and reliable prior, compelling the agent to rely on genuine visual grounding rather than superficial textual similarities. Extensive evaluations on FutureBench and NEPBench datasets demonstrate that our method achieves state-of-the-art performance, significantly outperforming larger MLLMs and validating the empirical effectiveness of our agentic, future-oriented reasoning paradigm.
Figures & tables
Figure 1: Key insight of VepAgent. (a) Vanilla VLM reasoning weakly models the future transition and predicts a coarse continuation of the observed action. (b) Option-driven reasoning is distracted by textual priors and prematurely predicts the cooking stage. (c) Causal-gap reasoning captures the ongoing buttering process, but misses the fine-grained visual state and selects the wrong continuation. (d) With tool-augmented visual evidence, VepAgent identifies the remaining unbuttered slice, correctly reasons over the unfinished action, and predicts the next event.
Figure 2: Pipeline of VepAgent. (i) Construct tool-augmented causal-transition CoT data by explicitly modeling the progression from terminal observed states to future events. (ii) Supervised fine-tuning initializes tool-use and causal-gap reasoning behaviors. (iii) GRPO optimizes the policy with accuracy, causal-gap, and anti-prior rewards. (iv) The trained agent dynamically integrates diagnostic tools for grounded future-event prediction.
Figure 3: Diagnostic tool library of VepAgent. FRE retrieves dense visual evidence from selected temporal windows, while RFM crops and magnifies task-relevant spatial regions. STT summarizes observed state trajectories and the terminal state, whereas VDE extracts fine-grained visible facts as textual evidence. All tools operate exclusively on the observed portion of the video.
Method
Frames
1-Hop
2-Hop
3-Hop
Interp.
AVG
GLM-4.1V-9B
32
29.90
41.90
52.20
47.30
44.38
LLaVA-NeXT-Video
32
48.80
49.30
40.00
44.40
45.24
Kimi-VL-A3B
32
44.30
42.80
51.30
51.90
48.87
MiMo-VL-7B
32
59.00
59.60
50.50
43.80
50.45
InternVL3-8B
32
54.30
58.00
63.20
54.40
56.72
Qwen2.5-VL-32B
32
66.50
62.70
63.20
55.20
59.94
Table 1: Event Prediction Acc. (%) on FutureBench. best and second best results are highlighted.
Method
Overall
Charades
VidSitu
YouCook2
GLM-4.1V-9B
43.73
45.98
28.18
64.48
InternVL3-8B
43.60
37.01
40.55
52.26
Kimi-VL-A3B
41.20
28.66
53.06
32.44
LLaVA-Video-7B
29.60
25.67
24.59
39.32
MiMo-VL-7B
44.53
40.47
40.83
52.46
Qwen2.5-VL-72B
47.50
35.59
41.05
64.48
Table 2: Event Prediction Acc. (%) on NEPBench. best and second best results are highlighted.
Figure 4: Case study between Qwen3-VL-4B and VepAgent-4B. The baseline predicts an already observed meat-addition event. StateTransitionTracker reconstructs the observed progression from meat addition to mixing and active cooking. Grounded in evidence, VepAgent identifies the unfinished cooking process, reasons over subsequent state evolution, and predicts the correct future event.
Method
1-Hop ↑
3-Hop ↑
AVG ↑
Tools
1-Hop ↑
3-Hop ↑
AVG ↑
(a) Qwen3-VL-4B
54.91
59.20
59.09
(e) w/o. STT
82.08
71.64
75.66
(b) 4B + SFT only
77.46
75.62
71.59
(f) w/o. FRE
82.08
75.62
76.99
(c) 4B + SFT + RL
79.77
71.14
75.09
(g) w/o. RFM
82.66
76.12
78.31
(d) 4B + SFT + RL + Tools ( G=4 )
83.82
76.12
78.79
(h) w/o. VDE
83.24
75.12
77.46
Reward Term
1-Hop ↑
3-Hop ↑
AVG ↑
Num of Gen.
1-Hop ↑
3-Hop ↑
AVG ↑
(i) R = Racc
78.61
76.12
74.62
(m) Num = 2
84.39
77.61
78.60
Table 3: Ablation studies (%) on FutureBench. Component-wise ablations are conducted with G=4 . best and second best results are highlighted.
Appendix figures & tables5 assets
Supplementary material from the paper’s appendix.
Appendix
Strategy
1-Hop
3-Hop
AVG
Avg Calls
(a) No Tools
79.77
71.14
75.09
0.00
(b) All Tools (Static)
76.30
70.65
74.91
4.00
(c) Random Tools
82.08
77.61
76.99
1.85
(d) Dynamic (Ours)
83.82
76.12
78.79
1.13
Appendix
Table 4: Ablation on tool selection strategies.
Tool Configuration
1-Hop
3-Hop
AVG
(a) Temporal Only (STT + FRE)
82.08
75.62
77.84
(b) Spatial Only (RFM + VDE)
80.92
71.64
75.38
(c) All (Ours)
83.82
76.12
78.79
Appendix
Table 5: Ablation on tool synergy between temporal and spatial dimensions.
Figure 5: Additional case study with FocusedEvidenceRetriever. The baseline repeatedly predicts the earlier fork-pressing action. FRE retrieves dense late-window visual evidence showing that the fork interaction has ended while the flatbread continues browning. VepAgent uses this temporal evidence to reason toward the subsequent cooking transition.
Figure 6: Additional case study with VisualDetailExtractor. The baseline repeats the already observed onion-addition event. VDE verifies fine-grained visual states, including partially translucent onions being stirred in the pot. VepAgent therefore recognizes that onion addition has completed and reasons toward the next ingredient stage.
Tool
Input
Output
Failure Handling
State Transition Tracker (STT)
observed video
textual event chain & final state
fallback to global video caption
Focused Evidence Retriever (FRE)
video & temporal window
dense frame sequence (contact sheet)
ignored if temporal window is invalid
Region Focus Magnifier (RFM)
video & predicted spatial boxes
magnified image patch
ignored if boxes are invalid
Visual Detail Extractor (VDE)
final frames & scene query
textual fine-grained facts
fallback to original query
Appendix
Table 6: Implementation details of the diagnostic tool library.
Video event prediction (VEP) requires models to infer unobserved future states from partial video evidence. Existing video MLLMs usually verbalize intermediate future reasoning in text space: once visual evidence is verbalized, fine-grained motion, geometry, and interaction cues can be lost, leading to plausible but visually ungrounded hallucinations. We introduce Future-L1, an interleaved latent visual reasoning framework that lets an MLLM alternate between language tokens and continuous latent visual spans during autoregressive decoding. To train this capability, we construct Future-L1-50K by selecting examples where future visual hints help prediction and align latent states to future-frame embeddings, then further optimize sampled latent trajectories with LA-DAPO, a latent-aware RL objective with outcome-contrastive and temporal-diversity rewards. Future-L1 achieves new state-of-the-art results on both benchmarks: on FutureBench, it improves Qwen3-VL-8B from 61.0 to 85.4 and exceeds the previous best Video-CoE by 10.4 points; on TwiFF-Bench, it improves the average score from 2.44 to 3.04. These results suggest that future-oriented video reasoning benefits from preserving intermediate visual semantics in latent space rather than translating every reasoning step into text.
Tianxiang Jiang, Linquan Wu, Sheng Xia +5
University of Science and Technology of China · Shanghai AI Laboratory · City University of Hong Kong +4
Understanding causal event relationships and achieving fine-grained temporal grounding in videos remain challenging for vision-language models (VLMs). We propose TEMPURA (Temporal Event Masked Prediction and Understanding for Reasoning in Action), a two-stage training framework that enhances the video temporal understanding of VLMs. Inspired by infilling techniques in language modeling, TEMPURA first performs masked event prediction, learning to reconstruct missing events and generate step-by-step causal explanations from dense event annotations. It then learns video segmentation and dense captioning, decomposing videos into non-overlapping events with detailed, timestamp-aligned descriptions. We train TEMPURA on VER, our large-scale dataset of 500K videos annotated with temporally aligned event descriptions and structured reasoning steps. Experiments on video temporal grounding and highlight detection benchmarks show that TEMPURA substantially improves strong base VLMs across model families and scales, confirming that combining event-level reasoning with fine-grained temporal segmentation is an effective recipe for video temporal understanding.
Jen-Hao Cheng, Yi-Hao Peng, Huapeng Zhou +11
University of Washington · Carnegie Mellon University · National Yang Ming Chiao Tung University +1
Accurately predicting future events is fundamental to content understanding and decision-making across various domains. While prior research has primarily focused on text or short-video scenarios, long-video event prediction, characterized by vast multimodal context and more complex narratives, remains underexplored. Meanwhile, although recent Long-Video Language Models (LVLMs), built on Large Language Models (LLMs) and Vision-Language Models (VLMs), have shown promise in long-video question answering and summarization, they struggle to generalize to event prediction, as they can neither precisely extract event-related details nor perform fine-grained analysis of event development. To address this gap, we propose VISTA, a multi-level event semantics mining framework for long-video event prediction. Initially, VISTA applies a character-centric visual prompt to precisely extract event-related visual details, enhancing detail-level semantics; subsequently, it employs a knowledge-enhanced iterative retrieval strategy, guiding the LLM to progressively construct logically coherent event chains, thereby improving event-level narratives; ultimately, VISTA adopts a human-like propose-then-retrieve strategy to generate diverse future-oriented proposals and integrate multi-level clues, producing robust and accurate predictions. Extensive experiments on real-world datasets validate the effectiveness of VISTA for long-video event prediction.