Current large video-language models (LVLMs) still face challenges when dealing with long videos, mainly because frames are often processed independently, making it difficult to capture temporal dependencies across events. Although retrieval-augmented approaches have been introduced to provide additional context, most of them operate at the frame or snippet level, which limits their ability to model how events evolve over time and relate to each other. In this paper, we propose Event Chain Retrieval-Augmented Generation (EC-RAG), a training-free framework that organizes video content into an explicit event chain before question answering. Instead of retrieving isolated frames or text segments, EC-RAG first partitions the video into semantically coherent segments, represents each segment using multi-modal signals, and then links them into a structured chain that preserves temporal order and captures inter-event relationships. Given a query, the system identifies relevant events within this chain and gathers supporting evidence from the associated modalities. Our approach offers several practical advantages: (i) event-level abstraction that better reflects how video content is naturally structured, enabling more reliable localization compared to frame-level retrieval; (ii) structured multi-modal fusion that aggregates speech, text, and visual cues at the event level, allowing complementary information to be more effectively utilized during reasoning; and (iii) plug-and-play compatibility with existing LVLM backbones, requiring no additional training or reliance on proprietary models. Experiments on Video-MME, MLVU, and LongVideoBench show that this event-centric design consistently outperforms frame-level retrieval baselines, highlighting the importance of modeling temporal structure for long-video understanding.
Figures & tables
Figure 1. Advantages of our EC-RAG. EC-RAG provides an event-centric, training-free pipeline with structured temporal understanding that is easily compatible with any LVLM.
Figure 2. Overview of the EC-RAG framework. The pipeline comprises four stages: (i) Query Decoupling parses the user question into modality-specific retrieval requests for ASR, OCR, and DET; (ii) CLIP-Variance Guided Multimodal Extraction selects informative keyframes through weighted scoring of visual similarity and frame variance, then extracts speech transcripts, on-screen text, and scene graphs via open-vocabulary object detection; (iii) Temporal Event Chain Construction segments the video into 30-second units and builds a chain of semantically coherent events, where each event integrates its associated multimodal evidence with explicit temporal relations; (iv) Evidence-augmented Multimodal Reasoning performs query-guided event localization to retrieve relevant events, then feeds the event chain summary, located events, and raw evidence to the LVLM for answer generation. By organizing video content as structured event sequences, EC-RAG captures temporal dynamics and inter-event dependencies for complex video understanding.
Model
#Text
LLM Params
Frames
Short
Medium
Long
Overall
Gain
Proprietary LVLMs
GPT-4o ( OpenAI, 2024 )
–
–
384
80.0
70.3
65.3
71.9
–
Gemini-1.5-Pro ( Team et al., 2024 )
–
–
0.5 fps
81.7
74.3
67.4
75.0
–
Open-Source LVLMs
Video-LLaVA ( Lin et al., 2024a )
–
7B
8
44.6
38.3
35.8
39.6
–
Video-LLaVA + EC-RAG
2.0K
7B
8
50.1
44.6
43.4
46.0
+6.4
Table 1. Performance on the Video-MME ( Fu et al., 2025 ) benchmark. #Text indicates the volume of retrieved textual context. EC-RAG is integrated into four open-source LVLMs spanning 8–32 input frames.
Model
#Params
Frames
Overall
Proprietary LVLMs
GPT-4o ( OpenAI, 2024 )
–
0.5 fps
64.6
Open-Source LVLMs
Video-CCAM ( Fei et al., 2024 )
14B
96
63.1
Video-XL ( Shu et al., 2025 )
7B
256
64.9
Aria ( Li et al., 2024 )
25.3B
256
70.6
Table 2. Overall accuracy on the multiple-choice split of the MLVU ( Zhou et al., 2024 ) benchmark. EC-RAG achieves the best result among all 7B-scale models.
Figure 3. Left: Grad-CAM heatmaps on query-relevant and query-irrelevant frames for the baseline and EC-RAG. Right: t-SNE projection of query, vision, and text features, showing tighter cross-modal alignment when EC-RAG is applied. Grad-CAM attention heatmaps and t-SNE feature projections comparing baseline and EC-RAG cross-modal alignment.
Model
#Params
Frames
Overall
VideoChat2-Mistral ( Li et al., 2025 )
7B
8
39.3
ShareGPT4Video ( Chen et al., 2024b )
7B
8
39.7
LLaVA-Next-Mistral ( Zhang et al., 2024d )
7B
8
49.1
PLLaVA ( Xu et al., 2024 )
34B
16
53.2
LLaVA-Video ( Zhang et al., 2024e )
7B
64
56.6
LLaVA-Video + Video-RAG ( Luo et al., 2024 )
7B
64
58.7
Table 3. Performance on the LongVideoBench ( Wu et al., 2024 ) validation set.
Figure 4. Qualitative comparison between the baseline LLaVA-Video and EC-RAG on a Video-MME example. The baseline confuses temporally distinct segments, whereas EC-RAG localizes the query-relevant event through its event chain and aggregates ASR and DET evidence to arrive at the correct answer. A qualitative case study showing how EC-RAG correctly answers a temporal reasoning question by localizing the relevant event and aggregating multimodal evidence.
ASR
OCR
DET
EC
CV
Short
Medium
Long
Overall
×
✓
✓
✓
✓
74.8
60.3
54.7
63.3
✓
×
✓
✓
✓
76.5
62.8
59.9
66.4
✓
✓
×
✓
✓
77.1
62.9
59.2
66.4
✓
✓
✓
×
✓
77.0
61.6
58.4
65.7
✓
✓
✓
✓
×
76.4
63.5
58.9
66.3
✓
✓
✓
✓
✓
77.8
64.3
60.4
67.5
Table 4. Component ablation of EC-RAG on Video-MME with LLaVA-Video-7B.
Appendix figures & tables6 assets
Supplementary material from the paper’s appendix.
Appendix
τ
#Token
Time
Short
Medium
Long
Overall
0.0
4.4K
52s
72.4
61.6
60.0
64.7
0.1
4.0K
46s
71.8
60.5
59.9
64.1
0.2
3.5K
23s
70.3
60.4
57.8
62.8
0.3
2.4K
18s
73.4
60.7
61.2
65.1
0.4
1.6K
13s
70.1
60.6
57.5
62.7
0.5
1.2K
12s
68.9
58.2
55.3
60.8
Appendix
Table 5. Performance with different retrieval thresholds on Video-MME ( Fu et al., 2025 ) with LLaVA-Video-7B ( Zhang et al., 2024e ) .
Model
Params
Frames
Short
Medium
Long
Overall
X=Video-LLaVA ( Lin et al., 2024a )
7B
8
44.6
38.3
35.8
39.6
X+Video-RAG ( Luo et al., 2024 )
7B
8
49.5
43.0
42.5
45.0
X+TV-RAG ( Cao et al., 2025 )
7B
8
49.7
44.3
42.6
45.5
X+EC-RAG
7B
8
50.1
44.6
43.4
46.0
X=LongVA ( Zhang et al., 2024b )
7B
32
60.9
49.3
44.0
51.4
X+Video-RAG ( Luo et al., 2024 )
7B
32
65.4
59.1
55.7
60.1
Appendix
Table 6. Comparison with other RAG methods across different backbones on Video-MME ( Fu et al., 2025 ) .
Figure 5. Accuracy of LLaVA-Video-7B ( Zhang et al., 2024e ) with and without EC-RAG under different frame budgets on Video-MME ( Fu et al., 2025 ) . EC-RAG provides consistent gains at all budgets, with the largest improvement at 16 frames. Bar chart comparing baseline and EC-RAG accuracy across 8, 16, and 32 frames on short, medium, and long videos.
Method
Frames
TmpR
ActR
ActRec
AttrP
SpaP
SpaR
TmpP
InfoS
OCR
ObjR
ObjRec
Count
AVG
Gain
LLaVA-Video ( Zhang et al., 2024e )
8
41.8
47.4
55.6
69.4
59.3
71.4
56.4
70.0
48.9
51.1
58.5
37.7
54.6
–
LLaVA-Video + EC-RAG
8
46.9
49.8
55.9
73.9
59.6
80.4
60.0
79.6
56.8
60.6
63.6
40.3
60.2
+5.6
LLaVA-Video ( Zhang et al., 2024e )
16
45.2
51.2
57.5
68.0
57.4
73.2
69.1
74.0
50.4
55.4
64.1
41.8
58.0
–
LLaVA-Video + EC-RAG
16
51.4
64.9
62.6
78.4
61.1
80.4
69.6
83.9
59.0
66.1
67.2
42.3
65.8
+7.8
Appendix
Table 7. Per-task-type accuracy (%) on Video-MME ( Fu et al., 2025 ) with LLaVA-Video-7B ( Zhang et al., 2024e ) as the backbone. Gain denotes the absolute improvement over the baseline. Best results per frame setting are in bold .
Figure 6. Qualitative example 1: a history documentary where the question requires reasoning about the temporal order of events after a key battle. The baseline confuses visually similar city depictions, while EC-RAG traces the event chain through ASR narration to identify the correct subsequent event. Qualitative comparison between baseline and EC-RAG on a history documentary temporal reasoning question.
Figure 7. Qualitative example 2: a science documentary where the question asks about the introduction order of four topics. The baseline is misled by recurring visual elements across segments, while EC-RAG recovers the correct topic sequence from the structured event chain. Qualitative comparison between baseline and EC-RAG on a science documentary topic ordering question.
Tianjin Key Lab of Intelligent Unmanned Swarm Tech & System Tianjin University Tianjin, China · Institute of Computing and Intelligence Harbin Institute of Technology (Shenzhen) Shenzhen, China · Tianjin Artificial Intelligence Innovation Center Tianjin, China +2
Department of Computer Science and Technology, Tsinghua University · Beijing National Research Center for Information Science and Technology, Tsinghua University