cs.CVOct 6, 2026

Decide Before You Look: Learning Which Retrieved Memories Deserve Pixels

Authors: Youxing LI

Organizations: The Chinese University of Hong Kong

Abstract

Multimodal assistants answer questions from long-term memories that contain images. After retrieval, each retrieved image reaches the answering model either as pixels, at about a thousand visual tokens per image, or as a stored text proxy that often misses the detail the question asks about. We find that the benefit of pixels usually comes from one or two retrieved memories, and that it can be predicted before the answering model runs, without reading any full-resolution image. In PixelTriage, a plug-in placed after retrieval, a small model that does not generate text reads the dialogue, a short note and a thumbnail of each retrieved memory and predicts how much its pixels would add. It is trained on synthetic memory episodes labeled by a frozen 27B model that answers each question with and without each memory's pixels. With a 7B answering model, PixelTriage lies on the accuracy--cost frontier of M3^3Exam, DMV and MemEye and uses 11--23% of the visual tokens without a significant loss of accuracy. On DMV it answers 2.9 times faster than opening all images. It outperforms retrieval order and uniform down-sizing at equal budgets and transfers to other memory systems and to a 397B answering model.

Figures & tables

Appendix figures & tables5 assets

Supplementary material from the paper’s appendix.

Appendix

Explore similar work

CardsList
  1. One Token per Multimodal Evidence: Latent Memory for Resource-Constrained QA

    Jun 9, 2026Zhi Zheng, Ziqiao Meng, Hao Luan +2Multimodal RAGMultimodal Memory

  2. PMMC: Prospective Multimodal Memory Compilation for Long-Term LVLM Agents

    Aug 2, 2026Jingyu Sun, Yan Lin, Yuyang Xue +10Memory-Augmented VLMsMultimodal Memory

  3. When Prompts Become Pixels: Prompt-Region Grounding for Multimodal Reasoning

    Aug 5, 2026Yongxin Wang, Ruizhe Zhou, Yueling Tang +4Multimodal QAMultimodal Reasoning