Hour-scale audio-visual question answering is constrained by a context dilemma: dense whole-recording encoding rapidly exhausts context limits, whereas uniform temporal compression severely dilutes fine-grained acoustic and visual evidence. We introduce LEAP, a framework where the model retrieves its own evidence without placing the whole recording in one context. LEAP divides a recording into fixed-duration blocks, applying a lightweight localization pass to each block to score short candidate windows. The highest-ranked windows are pooled and re-encoded in a single bounded answer pass. Consequently, the answer input and peak context remain independent of the recording duration. By decoupling evidence localization from reasoning, our framework can localize candidate temporal windows over pre-computed transcripts without decoding media frames, while preserving fine-grained visual and non-speech evidence by routing the final answering pass over raw audio-visual streams. LEAP trains both stages: a localization LoRA improves the selected windows, and an answer LoRA improves the answers read from the same windows. The block grid natively supports causal queries, enabling LEAP to support streaming inference without streaming-specific training. Across several AVQA benchmarks, LEAP improves over the Qwen3-Omni-30B-A3B baseline by 4.5-16.8%, and transfers to a second omni-modal backbone, MiniCPM-o 4.5, surpassing its published results by 3.1-13.0%.
Figures & tables
Figure 1: Overview of LEAP. Top: each recording is split into fixed-length blocks and transcribed once. Bottom: per question, a localization pass scores the candidate windows of every block (faint cells), from either the transcript with the base selector or the media with the localization LoRA; one answer pass re-reads the windows retained by block ranking (solid cells) from the recording.
System
TraceAV
LVOmni
VideoOdyssey (1–4h)
MMOU
gemini-3-flash-preview ∘
62.3
59.0
44.3
—
Qwen3.5-Omni-Plus ∘
—
—
43.0
—
Ming-Flash-Omni-2.0
51.7
34.6
—
—
Caption cascade into Qwen3-235B
—
—
—
47.9
Qwen3-Omni-30B, as published *
48.4
35.8
28.7
54.1
Qwen3-Omni-30B, with ASR transcript *
—
42.2
—
—
Table 1: Main results on Qwen3-Omni-30B. Accuracy (%). Upper block: closed-source systems ( ∘ ) as published. Middle block: open systems as published, then our backbone; * = the published number for our backbone; caption cascade = audio and visual captions read by a text-only model; ‡ = the benchmark’s official input recipe re-run on our stack, or an official-style whole-clip run where none is runnable. Lower block: LEAP without retrieval = LEAP’s answer LoRA reading the whole clip in one pass (on VideoOdyssey, its audio-montage recipe); LEAP without answer training = LEAP’s localization LoRA also answering the selected windows; — = no number. TraceAV = the mean over its twelve general sub-tasks. Bold = best, underline = second best in each column.
Figure 2: Across benchmarks and backbones. Solid = LEAP, dashed = baseline: (a) the same backbone on the whole recording (the causal prefix on StreamArena); (b) the backbone as published (official-style on our stack for VideoOdyssey and LVBench, native streaming on StreamArena).
Figure 3: The retrieval stage ablated. no media : stem and options only. (a) Accuracy; the localization adapter’s whisker (paired, vs. base selector) clears the base-selector bar where significant. (b) Evidence coverage: full bar = in a retained block, saturated = in an answer-pass window; OmniVideoBench on its timestamped subset, MMOU where both selectors stored a selection.
System
TraceAV
LVOmni
VideoOdyssey (1–4h)
MMOU
Media channel
63.6
45.8
53.7
66.3
Transcript channel
64.6
46.5
52.3
66.2
Localization tokens per question, media channel ÷ transcript channel
Media ÷ transcript
5.41 ×
4.74 ×
6.21 ×
4.93 ×
Query-time seconds per question (median)
Media channel
12.4
15.6
30.0
5.5
Table 2: Transcript-based retrieval against the media channel. Accuracy (%). Both channels feed the deployed answer pass. Localization tokens count the localization passes only; seconds = question-to-answer wall clock on one idle GPU over a duration-stratified sample, transcription charged to neither. TraceAV = accuracy over all its questions, hallucination sub-tasks included.
Figure 4: Accuracy vs. video length. All panels share one grid of length buckets, and each shows the buckets it populates. Error bars are each line’s video-clustered 95% interval.
evidence distance (min)
System
HR avg
≤ 5
5–15
15–30
> 30
Whole prefix
25.4
22.1
34.2
25.5
12.6
Uniform windows
20.2
23.2
25.8
17.9
9.3
LEAP – media
32.0
26.5
41.2
29.3
23.6
LEAP – transcript
30.1
24.3
36.1
30.4
24.7
Table 3: Historical retrospection under causal access on StreamArena, Qwen3-Omni-30B. Accuracy (%). Bins = minutes from the query time back to the evidence. Whole prefix reads everything before the query time in one pass; Uniform windows fills the same context with windows spread equally over it; LEAP – media and LEAP – transcript select the windows from the media blocks or from the transcript index built as the stream arrives, and answer them from the raw media. Bold = best, underline = second best in each column.
Appendix figures & tables29 assets
Supplementary material from the paper’s appendix.
Appendix
Backbone and adapters
Backbone
Qwen3-Omni-30B-A3B-Instruct, frozen in both stages
LoRA rank / scaling
r=16 / α=32 , identical for both adapters
Attachment points
attention query, key, value and output projections
Composition
sequential, one adapter resident at a time, never stacked
Training hardware
H100, bf16 throughout
Localization adapter ϕloc
Appendix
Table 4: Configuration of the two adapters.
MiniCPM-o 4.5
Backbone
MiniCPM-o 4.5
Attachment points
attention query, key, value and output projections, and the MLP gate, up and down projections
Localization objective
overlap-fraction BCE over up to eight 75 s candidate windows
Localization supervision grid
one level for every clip
Localization items per step
12
Data-parallel ranks
6 for the localization adapter, 3 for the answer adapter
Appendix
Table 5: MiniCPM-o 4.5 and the causal-access protocol.
Baseline
Frames
Audio
Read in
Official recipe ‡
TraceAV
up to 256
full track
Table 1 , Figure 2 a
LVOmniBench
128 at 3362
full track
Table 1 , Figure 2 a
VideoOdyssey
64
official montage
Tables 1 and 20 , Figure 2 a
MMOU
32
full track
Table 1 , Figure 2 a
Frame-matched whole clip
288
full track
Tables 1 and 7 , Appendix B.4
Appendix
Table 6: The single-pass baselines on Qwen3-Omni-30B. Each reads the whole recording, or on StreamArena the causal prefix, in one answer pass; uniform windows instead keeps the answer pass of LEAP and places its windows at equal spacing. Frames are spread evenly over what the baseline reads; full track = the native waveform of that span; official montage = 64 ten-second audio slices at evenly spaced starts, concatenated into one track, as cut by the benchmark’s released evaluation code. The official recipes of TraceAV and LVOmniBench also carry the benchmark’s own prompt.
System
TraceAV
TraceAV
LVOmni
VideoOdyssey (1–4h)
MMOU
General
Hallucination
Qwen3-Omni-30B
53.6
68.0
40.4
37.9
57.8
+ OmniVideo-100K SFT
52.4
65.0
39.3
43.1
70.1
+ OmniRAG-Agent loop
39.5
63.4
27.5
26.6
34.8
Qwen3-Omni-30B, as published *
48.4
67.5
—
—
—
+ AVP *
47.1
69.8
—
—
—
Appendix
Table 7: Concurrent systems on the same backbone. Accuracy (%). The first two rows read the same input: the whole recording sampled uniformly and frame-matched to the most frames our answer pass reads, with its full audio, or on VideoOdyssey the context-filled audio montage. OmniVideo-100K SFT = the released checkpoint fine-tuned on that instruction set, no retrieval stage ( Cai et al., 2026 ) . OmniRAG-Agent loop = the unmodified backbone inside that paper’s released retrieval agent, scored by our answer extractor ( Zhu et al., 2026 ) . * = the benchmark’s own runs, as TraceAV reports them ( Feng et al., 2026 ) : the bare backbone, and two untrained agents driving it, AVP ( Wang et al., 2026b ) and Audio-Guided OmniAgent ( Tao et al., 2025 ) . — = no number. TraceAV columns follow its two published tables: subtask macro-averages over the general and over the hallucination sub-tasks. Bold = best, underline = second best in each column.
Full denominator
Both forms fit
Benchmark
Whole clip
Selected windows
Whole clip
Selected windows
TraceAV
57.3
60.4
59.9
60.6
LVOmni
42.9
47.5
43.7
47.3
OmniVideoBench
42.3
43.3
42.3
43.3
MMOU
61.8
65.2
61.8
65.2
Appendix
Table 8: Input-form control. LEAP and the whole clip carry the same answer LoRA and no transcript outline, and differ in input form: the whole clip is read up to the backbone’s position limit, the selected windows at the answer-pass cap. Under full denominator a clip too long to encode whole inside the backbone’s position limit is scored wrong for the whole clip; both forms fit drops those questions. TraceAV = accuracy over all its questions, hallucination sub-tasks included. VideoOdyssey is absent: most of its recordings are too long to encode whole.
evidence distance (min)
L1
L2
L3
L4
System
Backbone
Query-time input
HR avg
≤ 5
5–15
15–30
> 30
Published, benchmark judge
StreamMind
Qwen3.5-397B-A17B
continuous memory
34.9
31.5
46.7
34.6
17.1
Qwen3.5-Omni ∘
—
whole prefix
35.8
31.5
44.8
34.1
25.4
AURA a
Qwen3-VL-8B
recent window
22.7
25.4
27.0
24.3
10.5
Appendix
Table 9: Historical retrospection under causal access on StreamArena. HR bins = minutes from the query time back to the evidence, under the benchmark’s own bin names. Query-time input = what the model reads when the question arrives. Published rows are the benchmark’s own table under its judge: ∘ = closed source, a = marked there as an author-finetuned backbone. LEAP – media scans the media blocks with the localization adapter mounted, Base selector the same blocks without it, and LEAP – transcript the transcript index built as the stream arrives; all three answer the retained windows from the raw media. Transcript only answers the windows of LEAP – transcript from their transcript instead. Bold = best, underline = second best in each column within a group.
Figure 5: LEAP in the streaming scenario. One time axis, cut by the query time; nothing to its right exists when the question is asked. Top : the raw archive and the transcript index built beside it. Bottom : the two selectors of LEAP and the one answer pass they share, which reads the retained windows from the archive and the outline from the index. Faint cells are candidate windows, solid cells the windows retained inside the blocks each selector keeps, and the outlined cell the slot held for the moment before the query, whether or not the ranking kept its block.
Figure 6: Query-to-answer time by stage on StreamArena. Seconds per historical-retrospection question on one H100, one group per pipeline stage: the localization passes over the causal blocks, decoding the retained windows from the archive, and the answer pass. The bar is the per-question median and the whisker above it reaches the 95th percentile.
System
TraceAV
LVOmni
VideoOdyssey
MMOU
LEAP
60.9
45.8
53.7
66.3
− transcript outline
57.8
47.5
51.2
65.2
− localization LoRA
56.2
43.5
46.7
62.1
− retrieval
54.8
42.9
41.9
61.8
Appendix
Table 10: Removing LEAP’s components. Accuracy (%) on Qwen3-Omni-30B. Each row removes one component from the row above , and the answer LoRA answers in every row. − transcript outline answers the same windows without the outline; − localization LoRA lets the bare backbone’s coarse scan pick the windows; − retrieval reads the whole clip in one pass (on VideoOdyssey, its audio-montage recipe). TraceAV = the mean over its twelve general sub-tasks. Bold = best, underline = second best in each column.
Accuracy (%)
Final-window coverage (%)
Ranking score
TraceAV
LVOmni
VideoOdyssey
MMOU
TraceAV
VideoOdyssey
MMOU
Base selector
58.3
39.0
41.6
56.2
92.0
57.3
82.7
Trial-answer confidence
58.8
39.0
41.5
56.3
90.2
48.3
82.7
SigLIP, frames
58.8
43.0
42.0
57.7
96.2
65.8
88.9
bge-m3, transcript
60.8
42.2
43.6
58.1
96.0
58.9
88.3
Localization adapter
59.2
43.3
45.6
61.1
96.2
71.7
94.8
Appendix
Table 11: An off-the-shelf retriever as the ranking score. Every row keeps the same candidate-window grid, block and window counts and answers with the localization adapter mounted; rows differ only in the score a candidate is ranked by. Under the trial-answer score the windows inside a block are still ranked by the base selector. All rows, the base selector and the localization adapter included, come from localization passes that split a recording’s partial tail block differently from the deployed ones, so they compare only with each other. TraceAV = accuracy over all its questions, hallucination sub-tasks included. Bold = best, underline = second best in each column.
Figure 7: Accuracy and evidence coverage against the retained-block count. Top : levels, accuracy on the left axis and evidence coverage on the right; triangles hold the deployed count and replace the block ranking by equally spaced blocks. Bottom : the paired accuracy difference against the deployed count, with its bootstrap interval; the deployed count is drawn open at zero. TraceAV runs only on the recordings that tile into more blocks than the deployed count keeps.
Recall
All
Ranking
Channel
Media decoded
Tokens / block
questions
decides
Content-free uniform
none
—
69.6
61.2
Listen-only channel
audio
6,855
71.7
67.0
Media channel
frames + audio
8,546
74.1
69.6
Transcript channel
none
1,579
74.3
70.0
Appendix
Table 12: What each channel’s selected windows retain on TraceAV. Content-free uniform spaces blocks and windows evenly and runs no pass. Media decoded = what a scoring pass decodes at query time; tokens / block = its measured length. Recall = the share of annotated evidence spans the selected windows deliver, over all questions and over those whose recording tiles into more blocks than the retained-block count ( ranking decides ). Bold = best, underline = second best in each recall column.
System
Transcript outline strategy
TraceAV
LVOmni
VideoOdyssey
MMOU
LEAP without answer training
Question-ranked
60.8
43.0
46.7
61.4
Uniform
59.8
42.8
45.0
61.0
None
58.2
42.7
43.6
61.1
LEAP
Question-ranked
63.6
45.8
53.7
66.3
Uniform
63.2
46.4
50.7
66.2
None
60.4
47.5
51.2
65.2
Appendix
Table 13: The transcript outline at the answer pass. Accuracy (%). Every row answers the same retrieved windows, grouped by system. Question-ranked (deployed) fills the outline question-first; Uniform thins every minute equally; None removes the outline. Bold = best, underline = second best within each group. TraceAV = accuracy over all its questions, hallucination sub-tasks included.
Group
No outline
+ transcript outline
MMOU: overlap with the annotated evidence interval
Evidence uncovered
48.0
55.9
Partially covered
63.8
65.5
Covered (at least half)
66.6
67.2
TraceAV: annotated question class
Spatiotemporal localization
30.8
50.7
Appendix
Table 14: Where the outline’s gain lands. Accuracy (%) of LEAP, with the transcript outline present or removed. MMOU questions are grouped by how much of their annotated evidence interval the retained windows overlap; TraceAV questions by their annotated class.
Figure 8: Accuracy against the answer-pass window count. The pooled windows are truncated to the top- k by window score. The deployed configuration keeps the whole pool, each curve’s last point.
Evidence coverage (%)
Partial block
Ranking score
Deletes
all questions
no partial block
retained (%)
gm=maxkσ(ℓm,k) (deployed)
—
70.2
78.7
67.6
maxksoftmax(ℓ)k
average, renormalized
63.8
73.8
62.5
maxkℓm,k−ℓˉm
average, subtracted
72.9
78.0
11.8
ℓˉm
margin
33.8
39.0
82.1
Km1∑kσ(ℓm,k)
margin; mean of σ
23.8
26.8
77.8
Appendix
Table 15: The block-ranking score, ablated by component. Each row re-ranks the same localization passes on VideoOdyssey. The deployed score is a monotone map of the block’s average logit plus its best window’s margin above it; Deletes = the component a variant removes, a dash neither. A partial block is a clip’s tail block, with fewer candidate windows; partial block retained = the share of questions retaining it. Bold = best, underline = second best in each coverage column.
Figure 9: Clue-grounded conversion on CG-Bench. LEAP against the whole clip; questions split by whether the selected windows overlap the human-labeled clue interval ( hit ) or not ( miss ). (a) Accuracy within each stratum. (b) Discordant flips, where exactly one of the two is correct.
Figure 10: Evidence-hit conversion inside the base-selector control. Top : each question flows from its final-window evidence status under the base selector (left) to its status under the localization LoRA (right); ribbon thickness is the share of questions on that transition, grey the unchanged. Bottom : the paired accuracy change, localization minus base, on the questions of each coloured transition, with its video-clustered interval. OmniVideoBench is read on its timestamped subset, MMOU on the questions with a stored selection under both selectors; LVOmniBench ships no timestamps.
Table 16: The audio ablation on TraceAV-Bench, at the two stages it can be read. (a) silences the localization pass alone and scores what the selected windows retain, with no answer pass; (b) keeps the audio-on selection’s windows, silences their answer pass (audio and transcript outline), and scores the answers.
Figure 11: The retrieval stage ablated on MiniCPM-o 4.5. Only the localization adapter differs. (a) Accuracy. (b) Evidence coverage; full bar = block coverage, saturated part = final-window coverage.
System
TraceAV
TraceAV
LVOmni
MMOU
General
Hallucination
test-15K
video-SALMONN-R 3 (8B) a
—
—
42.9
—
OmniAgent-RL-7B b
47.6
42.0
39.4
31.6
OmniAgent-RL-7B, as published *
55.0
46.1
—
—
MiniCPM-o 4.5, as published *
44.8
66.5
34.8
46.8
+ LEAP
57.8
69.6
44.4
51.0
Appendix
Table 17: The smaller backbone against its published rows and two systems of its size class. Accuracy (%). * = published number ( Feng et al., 2026 ; Tao et al., 2026b ; Goel et al., 2026 ) . a = a re-watching system on Qwen3-VL-8B with a Whisper encoder, as its paper reports it ( Li et al., 2026 ) . b = a native omni agent on Qwen2.5-Omni-7B, run by us under its own evaluator and step limit ( Xing et al., 2026 ) . — = no number. Columns follow each publication’s protocol: TraceAV’s two tables as subtask macro-averages, LVOmni micro, MMOU’s test-15K split under the benchmark’s own evaluator. Bold = best, underline = second best in each column.
Figure 12: MMOU by video duration: MiniCPM-o 4.5 on the test-15K split. Scored per duration bucket by the benchmark’s own evaluator. Dashed = the duration breakdown the benchmark paper reports for its own MiniCPM-o 4.5 run; solid = the same backbone carrying LEAP.
Figure 13: Accuracy vs. video length. Video-MME pools its two middle buckets. Muted = audio off at both passes, no transcript outline. Bars = video-clustered 95% interval, none on the muted line.
System
Certificate length (min)
Overall
[0, 0.5)
[0.5, 3)
[3, 15)
[15, 60)
[60, ∞ )
VideoOdyssey-V (vision-only)
Qwen3-Omni-30B base †
38.6
44.5
44.8
40.4
35.8
41.2
LEAP
44.0
49.5
41.3
40.4
43.1
44.1
VideoOdyssey-AV (audio-visual)
Qwen3-Omni-30B base † (+ audio montage)
34.0
33.7
43.0
35.3
38.9
36.7
Appendix
Table 18: Accuracy by certificate length on both VideoOdyssey tracks. Certificate length = the benchmark’s annotated length of the continuous span a viewer must watch to answer. † = the native-sampling whole clip on the same backbone, on the audio-visual track with the benchmark’s official audio montage. The LEAP rows re-answer their selected windows at 14 frames per window, matching the † frame count, and without the transcript outline. Bold = best per column and track.
Certificate length (min)
[0, 0.5)
[0.5, 3)
[3, 15)
[15, 60)
[60, ∞ )
Certificate hit (%)
LEAP
60.7
69.2
65.8
85.0
95.8
Random placement
9.4
20.5
43.0
88.7
100.0
Accuracy (%)
LEAP
55.7
49.6
53.3
37.6
61.1
Appendix
Table 19: The two readings of the certificate-length axis, tested on VideoOdyssey-AV. Certificate hit = an answered window overlaps the annotated certificate span. Random placement keeps the count of LEAP’s selected windows and redraws their positions over the recording. Bold = best in each column within a panel.
Figure 14: The annotated certificate span on VideoOdyssey-AV, by certificate length. Whole clip = the native-sampling whole clip. LEAP reads 14 frames per selected window. A moment-annotated question is absent from the certificate condition, the only one labelled with values.
Perception
Cognition
System
Count
ObRec
AcRec
VAR
AER
AAR
OCR
SFR
Cap
CaRea
EmRea
InRea
ObRea
SCR
SpRea
Order
Sum
TeGro
Overall
Human baseline
Human
75.0
85.0
85.4
81.0
78.7
71.1
81.1
87.9
80.6
81.0
74.4
79.3
73.7
75.0
82.6
71.0
70.8
93.2
80.7
Proprietary omni-modal LLMs
Gemini-2.5-Pro
25.9
41.3
40.0
44.8
38.4
44.6
45.8
50.3
74.5
50.0
46.8
48.1
42.9
53.2
23.3
38.0
60.0
37.7
43.9
Gemini-3-Flash
30.9
44.4
45.0
36.2
32.6
46.2
45.8
50.9
66.0
56.9
51.6
48.1
42.9
53.2
30.0
44.0
50.0
32.8
44.3
Appendix
Table 20: Per-task accuracy on the audio-visual track of VideoOdyssey. Columns are the benchmark’s task types; a question annotated with several is scored under each. Count counting, ObRec object, AcRec action, VAR visual attribute, AER acoustic event, AAR acoustic attribute and OCR character recognition, SFR speech fact retrieval, Cap captioning, CaRea causal, EmRea emotional, InRea intentional and ObRea object reasoning, SCR speech content reasoning, SpRea spatial reasoning, Order temporal ordering, Sum summarization, TeGro temporal grounding. ‡ = the benchmark’s official input recipe re-run on our stack; rows above it are published values. Bold = best, underline = second best in each column, the human row excluded.
Figure 15: Accuracy by montage audio coverage, VideoOdyssey-AV. Questions pool into coverage quantile bins, points at bin means, edges as minor ticks. Both systems answer with the answer LoRA in (a) and with the base weights in (b) . Bottom : paired difference.
Figure 16: Accuracy by evidence position, over the questions carrying evidence timestamps. The rows reduce a question’s annotated evidence span to one instant: its start, midpoint or end. LEAP and the whole clip carry the answer LoRA. Cells print bucket accuracy.
Weak supervision sets a practical regime for audio-visual sound source localization as dense spatial annotations are costly to obtain at scale. The task, however, remains challenging, as models must locate sound sources from temporally aligned audio-visual data without pixel-level supervision. Recent large-scale audio-visual retrieval models, trained at unprecedented scale, encode rich multimodal structure. We show their latent representations, though optimized for global alignment, can nonetheless enable fine-grained spatial grounding. While spatial detail is progressively lost in the upper layers of retrieval backbones due to global pooling, intermediate visual tokens retain highly structured spatial information. To exploit this, we introduce LAIP (\emph{Localization via Audio-Informed Pooling}), a framework that employs a lightweight \emph{Audio-informed Spatial Pooling} (AiSP) to replace the standard global aggregation module. By querying intermediate visual tokens with audio aligned at the frame level, LAIP recovers localized spatial information that is otherwise discarded by the retrieval pipeline, with the largest gains observed for PE-AV, a stack with underlying temporal aggregation. Our approach achieves state-of-the-art performance on AVSBench and AVATAR, nearly doubling previous results on the latter, improving average CIoU from 13.21 to 26.22.
Hugo Malard, Michel Olvera, Sanjeel Parekh +3
LTCI, Télécom Paris, Institut Polytechnique de Paris, France · Meta, Reality Labs Research · NVIDIA, France +1
Multimodal Large Language Models have achieved remarkable progress in short-form audio-video understanding, yet long-form audio-video comprehension remains challenged by limited context windows and severe information redundancy. To address these bottlenecks, we propose AVOC, a framework for long-form audio-video understanding in Omni-modal Large Language Models. AVOC introduces a learnable token compression module between the modality encoders and the LLM backbone. We reframe multimodal token compression as a top-K retrieval problem: given a fixed context budget, the module must retrieve a compact subset of tokens that best supports answering the user query. We draw inspiration from three classical Information Retrieval criteria for selecting informative units from a large candidate pool: relevance, importance, and diversity. AVOC instantiates each criterion as a tailored mechanism for audio-video understanding, and integrates them into a unified retrieval-style compression pipeline. Experiments show that AVOC achieves state-of-the-art performance on long-form audio-video benchmarks, surpassing the second-best model by 4.9 and 5.5 points in average accuracy on OmniVideoBench and LVOmniBench, respectively. Moreover, AVOC maintains robust performance on Audio-Video Needle-in-a-Haystack task at durations up to one hour.
Yijing Chen, Wenhui Tan, Xiaoyi Yu +7
1Gaoling School of Artificial Intelligence, Renmin University of China · 2Huawei Inc.
Long-video question answering requires identifying sparse yet critical evidence from videos containing thousands of frames under a constrained visual-token budget. Existing methods either select query-aware frames in a single pass or rely on timestamped text solely as retrieval guidance, leading to two key limitations. First, selected frames tend to cluster around local relevance peaks, and once the budget is exhausted, omitted evidence cannot be recovered. Second, textual and visual evidence remain weakly aligned. We propose GCR, a training-free framework that casts fixed-budget frame selection as a joint evidence curation problem. Ground converts timestamped text into temporal events, selects query-relevant real frame anchors, and renders each event text onto its temporally aligned frame. Cover supplements grounded events with direct visual anchors for complementary visual evidence and applies global maximal marginal relevance to preserve diverse context. Refine revisits omitted temporal regions and replaces the weakest revisable context frame with a real-frame medoid---but only when the medoid offers greater evidence value. GCR maintains a fixed number of chronologically ordered frames and requires no VLM training or architectural modification. Experiments on LongVideoBench and Video-MME, across three 7B backbones and frame budgets of 8, 32, and 64, demonstrate consistent improvements in long-video QA. With the 7B LLaVA-OV backbone and 32 frames, GCR achieves 64.25% and 62.15% on the two benchmarks, outperforming the strongest reproduced baselines by 2.54 and 1.93 percentage points, respectively.
Fan Wei, Siru Zhong, Runmin Dong +3
Department of Earth System Science, Tsinghua University · 5The Hong Kong University of Science and Technology (Guangzhou) · School of Artificial Intelligence, Sun Yat-Sen University +2