cs.CLSep 30, 2026

LEAP: Learned Block-wise Evidence Retrieval for Long Audio-Video Perception

Authors: Juyi Lin, Zhiqiang Lao, Jiali Cui, Lin Zhao, Pu Zhao, Dichang Zhang, Arman Akbari, Yu Qi, +4 more

Organizations: Northeastern University · Futurewei Technologies

Abstract

Hour-scale audio-visual question answering is constrained by a context dilemma: dense whole-recording encoding rapidly exhausts context limits, whereas uniform temporal compression severely dilutes fine-grained acoustic and visual evidence. We introduce LEAP, a framework where the model retrieves its own evidence without placing the whole recording in one context. LEAP divides a recording into fixed-duration blocks, applying a lightweight localization pass to each block to score short candidate windows. The highest-ranked windows are pooled and re-encoded in a single bounded answer pass. Consequently, the answer input and peak context remain independent of the recording duration. By decoupling evidence localization from reasoning, our framework can localize candidate temporal windows over pre-computed transcripts without decoding media frames, while preserving fine-grained visual and non-speech evidence by routing the final answering pass over raw audio-visual streams. LEAP trains both stages: a localization LoRA improves the selected windows, and an answer LoRA improves the answers read from the same windows. The block grid natively supports causal queries, enabling LEAP to support streaming inference without streaming-specific training. Across several AVQA benchmarks, LEAP improves over the Qwen3-Omni-30B-A3B baseline by 4.5-16.8%, and transfers to a second omni-modal backbone, MiniCPM-o 4.5, surpassing its published results by 3.1-13.0%.

Figures & tables

Appendix figures & tables29 assets

Supplementary material from the paper’s appendix.

Appendix

Explore similar work

CardsList
  1. Unlocking Spatial Grounding in Large Audio-Visual Retrieval models

    Jun 22, 2026Hugo Malard, Michel Olvera, Sanjeel Parekh +3Audio-Visual ReasoningSound Source Localization

  2. AVOC: Enhancing Hour-Level Audio-Video Understanding in Omni-Modal LLMs via Retrieval-Inspired Token Compression

    Jun 23, 2026Yijing Chen, Wenhui Tan, Xiaoyi Yu +7Audio-Visual ReasoningLarge Audio Language Models

  3. Ground, Cover, and Refine: Evidence-Centric Frame Selection for Long-Video Question Answering

    Aug 3, 2026Fan Wei, Siru Zhong, Runmin Dong +3Long Video Question AnsweringQuestion