Decide Before You Look: Learning Which Retrieved Memories Deserve Pixels
Organizations: The Chinese University of Hong Kong
Abstract
Multimodal assistants answer questions from long-term memories that contain images. After retrieval, each retrieved image reaches the answering model either as pixels, at about a thousand visual tokens per image, or as a stored text proxy that often misses the detail the question asks about. We find that the benefit of pixels usually comes from one or two retrieved memories, and that it can be predicted before the answering model runs, without reading any full-resolution image. In PixelTriage, a plug-in placed after retrieval, a small model that does not generate text reads the dialogue, a short note and a thumbnail of each retrieved memory and predicts how much its pixels would add. It is trained on synthetic memory episodes labeled by a frozen 27B model that answers each question with and without each memory's pixels. With a 7B answering model, PixelTriage lies on the accuracy--cost frontier of MExam, DMV and MemEye and uses 11--23% of the visual tokens without a significant loss of accuracy. On DMV it answers 2.9 times faster than opening all images. It outperforms retrieval order and uniform down-sizing at equal budgets and transfers to other memory systems and to a 397B answering model.
Figures & tables
| M 3 Exam | DMV | MemEye | |||||||
|---|---|---|---|---|---|---|---|---|---|
| Delivery policy | |||||||||
| Text proxies only | 67.58 | 54.45 | 37.47 | ||||||
| Open all retrieved images | 75.08 | 64.75 | 48.52 | ||||||
| Random | 68.96 | 71.06 | 72.36 | 55.90 | 57.85 | 59.60 | 40.03 | 41.64 | 42.59 |
| Retrieval order | 73.47 | 74.50 | 74.73 | 58.15 | 59.55 | 59.55 | 47.30 | 47.17 | 47.30 |
| General decision model | 69.30 | 72.13 | 73.47 | 57.00 | 57.40 | 58.85 | 43.94 | 46.77 | 47.98 |
| Setting | Benchmark | Text | Rank@1 | Ours@1 | All |
|---|---|---|---|---|---|
| UniversalRAG | DMV | 19.45 | 20.60 | 25.05 | 25.55 |
| UniversalRAG | MemEye | 32.61 | 39.08 | 42.72 | 41.91 |
| NaiveRAG | DMV | 55.05 | 62.90 | 66.65 | 68.10 |
| NaiveRAG | MemEye | 31.94 | 36.25 | 40.16 | 41.51 |
| Qwen3.5-397B | M 3 Exam | 67.2 | 77.1 | 78.6 | 80.0 |
| Qwen3.5-397B | DMV | 62.2 | 71.5 | 77.9 | 84.6 |
| Answering | Open all | Ours@1 | Ratio | |
| DMV time (s) | 7B | 2.05 | 0.71 | 2.9 |
| 27B | 2.77 | 1.07 | 2.6 | |
| M 3 Exam time (s) | 7B | 0.99 | 0.90 | 1.1 |
| 27B | 2.62 | 2.41 | 1.1 | |
| DMV TFLOP | 7B | 238 | 65 | 27% |
| 27B | 670 | 179 | 27% |
| System One input | AUROC | Top-1 is ref. (%) | Ref. helps |
|---|---|---|---|
| Memory headers | 0.718 .003 | 41.1 0.4 | 0.666 .005 |
| + text notes | 0.809 .005 | 75.4 0.8 | 0.782 .002 |
| + thumbnails | 0.835 .001 | 86.7 3.4 | 0.847 .005 |
Appendix figures & tables5 assets
Supplementary material from the paper’s appendix.
Appendix
| Ours@ minus | M 3 Exam | DMV | MemEye | |
|---|---|---|---|---|
| Open all retrieved images | 1 | 0.65 [ 0.34, 1.68] | 2.45 [ 4.55, 0.30] | 0.54 [ 3.10, 4.18] |
| 2 | 0.34 [ 0.54, 1.26] | 1.40 [ 3.25, 0.40] | 1.75 [ 1.48, 5.12] | |
| 3 | 0.38 [ 0.27, 1.11] | 1.30 [ 3.10, 0.45] | 2.02 [ 1.08, 5.12] | |
| Retrieval order | 1 | 2.26 [ 1.07, 3.52] | 4.15 [ 2.45, 5.90] | 1.75 [ 1.21, 4.72] |
| 2 | 0.92 [ 0.19, 2.06] | 3.80 [ 2.20, 5.45] | 3.10 [ 0.00, 6.33] | |
| 3 | 0.73 [ 0.19, 1.64] | 3.90 [ 2.30, 5.55] | 3.23 [ 0.40, 6.06] |
| Setting | Benchmark | Ours@1 Rank@1 | Ours@1 All |
|---|---|---|---|
| UniversalRAG | DMV | 4.45 [ 3.00, 5.95] | 0.50 |
| UniversalRAG | MemEye | 3.64 [ 0.40, 6.87] | 0.81 |
| NaiveRAG | DMV | 3.75 [ 2.05, 5.55] | 1.45 |
| NaiveRAG | MemEye | 3.91 [ 0.94, 7.01] | 1.35 |
| Qwen3.5-397B | M 3 Exam | 1.6 [ 0.2, 3.0] | 1.4 [ 2.5, 0.3] |
| Qwen3.5-397B | DMV | 6.5 [ 4.6, 8.3] | 6.7 [ 8.2, 5.3] |
| Rate | All test | Cross-fit | 50 unlabeled | Random | Cost | |
|---|---|---|---|---|---|---|
| M 3 Exam | 30% | 75.08 | 75.12 [75.00, 75.23] | 75.17 [74.38, 76.03] | 70.03 | 6.0% |
| 50% | 75.65 | 75.69 [75.65, 75.80] | 75.70 [75.04, 76.37] | 71.66 | 10.8% | |
| DMV | 30% | 57.15 | 57.15 [56.95, 57.40] | 57.25 [55.84, 58.53] | 56.80 | 3.0% |
| 50% | 58.65 | 58.72 [58.50, 58.95] | 58.71 [57.84, 59.63] | 58.38 | 5.0% | |
| MemEye | 30% | 41.78 | 41.99 [41.37, 43.00] | 42.18 [38.78, 45.64] | 40.94 | 3.4% |
| 50% | 45.69 | 45.54 [45.01, 46.09] | 45.64 [43.30, 47.98] | 43.26 | 5.6% |
| Opened | Cost | Gate | Random gate | All | Gate random | |
|---|---|---|---|---|---|---|
| M 3 Exam | 25% | 4.9% | 75.00 | 69.63 | 75.08 | 5.37 [ 3.90, 6.91] |
| DMV | 77% | 7.7% | 61.40 | 60.48 | 64.75 | 0.92 [ 0.08, 1.73] |
| MemEye | 83% | 9.4% | 47.57 | 47.12 | 48.52 | 0.45 [ 1.05, 1.85] |
| Captions | Text | All | Rank@1 | Ours@1 | Ours@2 | Ours@3 | |
|---|---|---|---|---|---|---|---|
| M 3 Exam | benchmark, replaced | 67.58 | 75.08 | 73.47 | 75.73 | 75.42 | 75.46 |
| our notes, replaced | 67.97 | 75.04 | 73.43 | 75.42 | 75.31 | 75.11 | |
| none | 67.01 | 75.00 | 73.39 | 75.69 | 75.84 | 75.80 | |
| DMV | benchmark, kept | 54.45 | 64.75 | 58.15 | 62.30 | 63.35 | 63.45 |
| benchmark, replaced | 54.40 | 60.00 | 53.70 | 53.10 | 56.35 | 57.20 | |
| our notes, replaced | 28.45 | 60.00 | 49.30 | 60.15 | 61.70 | 60.35 |