LEAP: Learned Block-wise Evidence Retrieval for Long Audio-Video Perception
Organizations: Northeastern University · Futurewei Technologies
Abstract
Hour-scale audio-visual question answering is constrained by a context dilemma: dense whole-recording encoding rapidly exhausts context limits, whereas uniform temporal compression severely dilutes fine-grained acoustic and visual evidence. We introduce LEAP, a framework where the model retrieves its own evidence without placing the whole recording in one context. LEAP divides a recording into fixed-duration blocks, applying a lightweight localization pass to each block to score short candidate windows. The highest-ranked windows are pooled and re-encoded in a single bounded answer pass. Consequently, the answer input and peak context remain independent of the recording duration. By decoupling evidence localization from reasoning, our framework can localize candidate temporal windows over pre-computed transcripts without decoding media frames, while preserving fine-grained visual and non-speech evidence by routing the final answering pass over raw audio-visual streams. LEAP trains both stages: a localization LoRA improves the selected windows, and an answer LoRA improves the answers read from the same windows. The block grid natively supports causal queries, enabling LEAP to support streaming inference without streaming-specific training. Across several AVQA benchmarks, LEAP improves over the Qwen3-Omni-30B-A3B baseline by 4.5-16.8%, and transfers to a second omni-modal backbone, MiniCPM-o 4.5, surpassing its published results by 3.1-13.0%.
Figures & tables
| System | TraceAV | LVOmni | VideoOdyssey (1–4h) | MMOU |
| gemini-3-flash-preview | 62.3 | 59.0 | 44.3 | — |
| Qwen3.5-Omni-Plus | — | — | 43.0 | — |
| Ming-Flash-Omni-2.0 | 51.7 | 34.6 | — | — |
| Caption cascade into Qwen3-235B | — | — | — | 47.9 |
| Qwen3-Omni-30B, as published * | 48.4 | 35.8 | 28.7 | 54.1 |
| Qwen3-Omni-30B, with ASR transcript * | — | 42.2 | — | — |
| System | TraceAV | LVOmni | VideoOdyssey (1–4h) | MMOU |
|---|---|---|---|---|
| Media channel | 63.6 | 45.8 | 53.7 | 66.3 |
| Transcript channel | 64.6 | 46.5 | 52.3 | 66.2 |
| Localization tokens per question, media channel transcript channel | ||||
| Media transcript | 5.41 | 4.74 | 6.21 | 4.93 |
| Query-time seconds per question (median) | ||||
| Media channel | 12.4 | 15.6 | 30.0 | 5.5 |
| evidence distance (min) | |||||
|---|---|---|---|---|---|
| System | HR avg | 5 | 5–15 | 15–30 | 30 |
| Whole prefix | 25.4 | 22.1 | 34.2 | 25.5 | 12.6 |
| Uniform windows | 20.2 | 23.2 | 25.8 | 17.9 | 9.3 |
| LEAP – media | 32.0 | 26.5 | 41.2 | 29.3 | 23.6 |
| LEAP – transcript | 30.1 | 24.3 | 36.1 | 30.4 | 24.7 |
Appendix figures & tables29 assets
Supplementary material from the paper’s appendix.
Appendix
| Backbone and adapters | |
|---|---|
| Backbone | Qwen3-Omni-30B-A3B-Instruct, frozen in both stages |
| LoRA rank / scaling | / , identical for both adapters |
| Attachment points | attention query, key, value and output projections |
| Composition | sequential, one adapter resident at a time, never stacked |
| Training hardware | H100, bf16 throughout |
| Localization adapter | |
| MiniCPM-o 4.5 | |
|---|---|
| Backbone | MiniCPM-o 4.5 |
| Attachment points | attention query, key, value and output projections, and the MLP gate, up and down projections |
| Localization objective | overlap-fraction BCE over up to eight s candidate windows |
| Localization supervision grid | one level for every clip |
| Localization items per step | |
| Data-parallel ranks | for the localization adapter, for the answer adapter |
| Baseline | Frames | Audio | Read in |
| Official recipe ‡ | |||
| TraceAV | up to | full track | Table 1 , Figure 2 a |
| LVOmniBench | at | full track | Table 1 , Figure 2 a |
| VideoOdyssey | official montage | Tables 1 and 20 , Figure 2 a | |
| MMOU | full track | Table 1 , Figure 2 a | |
| Frame-matched whole clip | full track | Tables 1 and 7 , Appendix B.4 | |
| System | TraceAV | TraceAV | LVOmni | VideoOdyssey (1–4h) | MMOU |
|---|---|---|---|---|---|
| General | Hallucination | ||||
| Qwen3-Omni-30B | 53.6 | 68.0 | 40.4 | 37.9 | 57.8 |
| + OmniVideo-100K SFT | 52.4 | 65.0 | 39.3 | 43.1 | 70.1 |
| + OmniRAG-Agent loop | 39.5 | 63.4 | 27.5 | 26.6 | 34.8 |
| Qwen3-Omni-30B, as published * | 48.4 | 67.5 | — | — | — |
| + AVP * | 47.1 | 69.8 | — | — | — |
| Full denominator | Both forms fit | |||
|---|---|---|---|---|
| Benchmark | Whole clip | Selected windows | Whole clip | Selected windows |
| TraceAV | 57.3 | 60.4 | 59.9 | 60.6 |
| LVOmni | 42.9 | 47.5 | 43.7 | 47.3 |
| OmniVideoBench | 42.3 | 43.3 | 42.3 | 43.3 |
| MMOU | 61.8 | 65.2 | 61.8 | 65.2 |
| evidence distance (min) | |||||||
| L1 | L2 | L3 | L4 | ||||
| System | Backbone | Query-time input | HR avg | 5 | 5–15 | 15–30 | 30 |
| Published, benchmark judge | |||||||
| StreamMind | Qwen3.5-397B-A17B | continuous memory | 34.9 | 31.5 | 46.7 | 34.6 | 17.1 |
| Qwen3.5-Omni | — | whole prefix | 35.8 | 31.5 | 44.8 | 34.1 | 25.4 |
| AURA a | Qwen3-VL-8B | recent window | 22.7 | 25.4 | 27.0 | 24.3 | 10.5 |
| System | TraceAV | LVOmni | VideoOdyssey | MMOU |
|---|---|---|---|---|
| LEAP | 60.9 | 45.8 | 53.7 | 66.3 |
| transcript outline | 57.8 | 47.5 | 51.2 | 65.2 |
| localization LoRA | 56.2 | 43.5 | 46.7 | 62.1 |
| retrieval | 54.8 | 42.9 | 41.9 | 61.8 |
| Accuracy (%) | Final-window coverage (%) | ||||||
|---|---|---|---|---|---|---|---|
| Ranking score | TraceAV | LVOmni | VideoOdyssey | MMOU | TraceAV | VideoOdyssey | MMOU |
| Base selector | 58.3 | 39.0 | 41.6 | 56.2 | 92.0 | 57.3 | 82.7 |
| Trial-answer confidence | 58.8 | 39.0 | 41.5 | 56.3 | 90.2 | 48.3 | 82.7 |
| SigLIP, frames | 58.8 | 43.0 | 42.0 | 57.7 | 96.2 | 65.8 | 88.9 |
| bge-m3, transcript | 60.8 | 42.2 | 43.6 | 58.1 | 96.0 | 58.9 | 88.3 |
| Localization adapter | 59.2 | 43.3 | 45.6 | 61.1 | 96.2 | 71.7 | 94.8 |
| Recall | ||||
| All | Ranking | |||
| Channel | Media decoded | Tokens / block | questions | decides |
| Content-free uniform | none | — | 69.6 | 61.2 |
| Listen-only channel | audio | 6,855 | 71.7 | 67.0 |
| Media channel | frames + audio | 8,546 | 74.1 | 69.6 |
| Transcript channel | none | 1,579 | 74.3 | 70.0 |
| System | Transcript outline strategy | TraceAV | LVOmni | VideoOdyssey | MMOU |
|---|---|---|---|---|---|
| LEAP without answer training | Question-ranked | 60.8 | 43.0 | 46.7 | 61.4 |
| Uniform | 59.8 | 42.8 | 45.0 | 61.0 | |
| None | 58.2 | 42.7 | 43.6 | 61.1 | |
| LEAP | Question-ranked | 63.6 | 45.8 | 53.7 | 66.3 |
| Uniform | 63.2 | 46.4 | 50.7 | 66.2 | |
| None | 60.4 | 47.5 | 51.2 | 65.2 |
| Group | No outline | + transcript outline |
|---|---|---|
| MMOU: overlap with the annotated evidence interval | ||
| Evidence uncovered | 48.0 | 55.9 |
| Partially covered | 63.8 | 65.5 |
| Covered (at least half) | 66.6 | 67.2 |
| TraceAV: annotated question class | ||
| Spatiotemporal localization | 30.8 | 50.7 |
| Evidence coverage (%) | Partial block | |||
|---|---|---|---|---|
| Ranking score | Deletes | all questions | no partial block | retained (%) |
| (deployed) | — | 70.2 | 78.7 | 67.6 |
| average, renormalized | 63.8 | 73.8 | 62.5 | |
| average, subtracted | 72.9 | 78.0 | 11.8 | |
| margin | 33.8 | 39.0 | 82.1 | |
| margin; mean of | 23.8 | 26.8 | 77.8 | |
| System | TraceAV | TraceAV | LVOmni | MMOU |
|---|---|---|---|---|
| General | Hallucination | test-15K | ||
| video-SALMONN-R 3 (8B) a | — | — | 42.9 | — |
| OmniAgent-RL-7B b | 47.6 | 42.0 | 39.4 | 31.6 |
| OmniAgent-RL-7B, as published * | 55.0 | 46.1 | — | — |
| MiniCPM-o 4.5, as published * | 44.8 | 66.5 | 34.8 | 46.8 |
| + LEAP | 57.8 | 69.6 | 44.4 | 51.0 |
| System | Certificate length (min) | Overall | ||||
| [0, 0.5) | [0.5, 3) | [3, 15) | [15, 60) | [60, ) | ||
| VideoOdyssey-V (vision-only) | ||||||
| Qwen3-Omni-30B base † | 38.6 | 44.5 | 44.8 | 40.4 | 35.8 | 41.2 |
| LEAP | 44.0 | 49.5 | 41.3 | 40.4 | 43.1 | 44.1 |
| VideoOdyssey-AV (audio-visual) | ||||||
| Qwen3-Omni-30B base † (+ audio montage) | 34.0 | 33.7 | 43.0 | 35.3 | 38.9 | 36.7 |
| Certificate length (min) | |||||
| [0, 0.5) | [0.5, 3) | [3, 15) | [15, 60) | [60, ) | |
| Certificate hit (%) | |||||
| LEAP | 60.7 | 69.2 | 65.8 | 85.0 | 95.8 |
| Random placement | 9.4 | 20.5 | 43.0 | 88.7 | 100.0 |
| Accuracy (%) | |||||
| LEAP | 55.7 | 49.6 | 53.3 | 37.6 | 61.1 |
| Perception | Cognition | ||||||||||||||||||
| System | Count | ObRec | AcRec | VAR | AER | AAR | OCR | SFR | Cap | CaRea | EmRea | InRea | ObRea | SCR | SpRea | Order | Sum | TeGro | Overall |
| Human baseline | |||||||||||||||||||
| Human | 75.0 | 85.0 | 85.4 | 81.0 | 78.7 | 71.1 | 81.1 | 87.9 | 80.6 | 81.0 | 74.4 | 79.3 | 73.7 | 75.0 | 82.6 | 71.0 | 70.8 | 93.2 | 80.7 |
| Proprietary omni-modal LLMs | |||||||||||||||||||
| Gemini-2.5-Pro | 25.9 | 41.3 | 40.0 | 44.8 | 38.4 | 44.6 | 45.8 | 50.3 | 74.5 | 50.0 | 46.8 | 48.1 | 42.9 | 53.2 | 23.3 | 38.0 | 60.0 | 37.7 | 43.9 |
| Gemini-3-Flash | 30.9 | 44.4 | 45.0 | 36.2 | 32.6 | 46.2 | 45.8 | 50.9 | 66.0 | 56.9 | 51.6 | 48.1 | 42.9 | 53.2 | 30.0 | 44.0 | 50.0 | 32.8 | 44.3 |