cs.CVSep 29, 2026

LazySloth: Bounded LLM-based Lazy Tree Search for Fast Long Video Comprehension

Authors: Arka Mukherjee, Kaleen Shrestha, Larissa Zhu, Maja Matarić

Organizations: School of Computer Engineering, Kalinga Institute of Industrial Technology (KIIT) Bhubaneswar · Computer Science Department, Viterbi School of Engineering, University of Southern California

Abstract

Modern vision-language models (VLMs) have shown promising results in long-video understanding due to the rich semantic information they can capture. However, most methods focus on coarse captioning of extracted image frames that are computationally inefficient and require models with large context windows. While past work has explored efficient methods through multimodal retrieval-augmented generation (RAG), they rely on lossy embeddings that lose temporal context and fine-grained detail. Few works to date have investigated how VLM-based query-relevant information retrieval can be optimized. We introduce LazySloth, an efficient tree-based search method that speeds up video comprehension and retrieval tasks 2.9-8.3x (compared to existing agentic methods) through bounded captioning of portions of the video considered irrelevant by a VLM of the video. Compared to contemporary specialized video-understanding VLMs and RAG-based methods, LazySloth achieved similar or better final task accuracy across two recent open-source base VLMs--Gemma 4 31B and Qwen3.6 27B--across four benchmarks. LazySloth reduced the gap between the base open-source model and a closed-source model, GPT-4o. Ablations showed that replacing VLM scene understanding with CLIP-based retrieval cost 8.8-19.9% in accuracy, while lazy tree construction matches eager construction at a fraction of the captioning cost. With LazySloth, we demonstrate the possibility of faster long-video comprehension without substantial loss in performance.

Figures & tables

Appendix figures & tables10 assets

Supplementary material from the paper’s appendix.

Appendix

Explore similar work

CardsList
  1. PREM: Prefix-Steered Recurrent Memory for Long-Video Understanding

    Sep 20, 2026Siru Zhong, Qiongyan Wang, Xiaohui Lv +5Frozen Vision-Language ModelsLong-Video Benchmarks

  2. Incentivizing Vision Language Models to Search for Long Video Question Answering

    Jul 3, 2026Harsh Goel, S P Sharan, Sahil Shah +4Long Video Question AnsweringVideo Understanding