VideoTapestry: Query-Adaptive Memory Refinement for Multi-Agent Long-Video Understanding
Organizations: University of Science and Technology of China · Jilin University · Hangzhou Dianzi University · Institute of Artificial Intelligence, Hefei Comprehensive National Science Center
Abstract
Long-video understanding places substantial demands on memory, as answering questions often requires retrieving information distributed across extended temporal spans. Existing approaches broadly follow two paradigms: query-driven exploration, which is sensitive to localization errors, and query-independent memory construction, which may omit question-specific details. We introduce VideoTapestry, a training-free multi-agent framework that adapts a preconstructed hierarchical video memory through coarse-to-fine, query-driven refinement. The preconstructed memory organizes video content into three levels, capturing global narrative context, event-level temporal structure, and fine-grained relational evidence, respectively. To support coarse-to-fine localization and observation, we assign a specialized agent to each level, keeping retrieval and refinement within a scale-specific context. Guided by the query, these agents revisit relevant video regions and enrich layer-wise memories with targeted multimodal observations. Their refinements are assembled according to the original hierarchy into a composite query-adaptive memory, preserving global context in a compact form while retaining fine-grained evidence along query-relevant branches for final reasoning. Compared with direct GPT-5.5 inference, VideoTapestry achieves absolute accuracy gains of 17.2%, 14.9%, 9.8%, and 7.0% on LVBench, LongVideoBench (Long), Video-MME (Long), and EgoSchema, respectively, achieving the state-of-the-art results among all competitors.
Figures & tables
| Method | Query-time Model(s) | LVBench | LongVideoBench | Video-MME(w/o) | EgoSchema | |
|---|---|---|---|---|---|---|
| Overall | Val | Long | Long | Val | ||
| Direct Proprietary VLMs | ||||||
| Gemini-2.5-Pro | Gemini-2.5-Pro | 72.0 | 71.0 | 68.6 | 75.9 | 72.8 |
| Gemini-3.1-Pro | Gemini-3.1-Pro | 78.2 | 78.6 | 77.0 | 80.3 | 76.4 |
| OpenAI o3 | OpenAI o3 | 57.1 | 66.7 | 60.6 | 64.7 | 63.2 |
| Direct Open-Source VLMs | ||||||
| Method | ER | EU | Rea | KIR | Sum | TG | Avg. |
|---|---|---|---|---|---|---|---|
| Direct VLMs | |||||||
| GPT-4o | 48.9 | 49.5 | 50.3 | 48.1 | 50.0 | 40.9 | 48.9 |
| OpenAI o3 | 57.6 | 56.4 | 50.8 | 62.9 | 67.2 | 46.8 | 57.1 |
| GPT-5.5 (Direct) | 63.2 | 66.0 | 72.7 | 63.8 | 65.4 | 43.1 | 64.3 |
| Structured and Agentic Methods | |||||||
| VideoTree | 30.3 | 25.1 | 31.9 | 26.5 | 25.5 | 27.7 | 28.8 |
| Backbone | Direct | VideoTapestry | Improve. |
|---|---|---|---|
| GPT-5.5 | 78.0 | 87.0 | |
| GPT-5 | 75.5 | 82.5 | |
| Qwen3.7-Plus | 75.0 | 83.0 | |
| Qwen3.6-27B | 65.5 | 75.0 |
Appendix figures & tables1 asset
Supplementary material from the paper’s appendix.