Long-horizon tasks require preserving and later recovering cross-session evidence under a bounded, query-blind memory budget. Existing compression can discard fine-grained visual cues or conflate semantically similar but incompatible observations. We present C3M, a cross-session multimodal memory organization that maintains a bounded active index over persistent source text-image evidence. Relation-aware updates consolidate safe redundancy while preserving complementary and incompatible records. At query time, budgeted routing selects useful index pages and expands their associated source evidence under a fixed reader budget. Together, these mechanisms establish a compact, provenance-preserving multimodal memory organization for cross-session long-horizon tasks, retaining temporal distinctions and source links required for reliable downstream reasoning. Code is available at https://github.com/HuzhouNLP/C3M.
Figures & tables
Figure 1: Query-blind multimodal memory challenges. (a) Similar records can be distinct. (b) Compression can lose visual details or source links. (c) Routing can leave stored evidence unrecovered.
Figure 2: Overview of C3M. (a) During query-blind memory maintenance, RAMU applies relation-aware Merge , Update , Add , or Keep actions to incoming multimodal sessions. Compact, provenance-linked entries are maintained in the active index under a fixed write budget, while original evidence is preserved in the persistent Cold Store. (b) Given a query, BMER selects index pages (P) and entries (E), then retrieves linked Raw Nodes and original images from the Cold Store under a joint reader budget over pages, nodes, images, and text for answer generation.
Model
Method
IE
MSR
TR
KU
AR
Overall
32K
64K
128K
32K
64K
128K
32K
64K
128K
32K
64K
128K
32K
64K
128K
32K
64K
128K
GPT-5.6 Sol
ReSum
44.26
37.70
42.62
22.86
20.00
22.86
64.58
56.25
43.75
34.48
24.14
17.24
59.09
50.00
45.45
45.64
38.46
35.90
MovieChat
73.77
65.57
57.38
45.71
42.86
54.29
70.83
70.83
58.33
44.83
48.28
44.83
86.36
86.36
72.73
65.13
62.56
56.92
MemRefine
67.21
60.66
65.57
65.71
57.14
57.14
72.92
75.00
70.83
51.72
51.72
44.83
86.36
81.82
63.64
68.21
64.62
62.05
CDRs
72.13
67.21
68.85
45.71
51.43
51.43
64.58
66.67
72.92
48.28
51.72
51.72
90.91
86.36
81.82
64.10
64.10
65.64
C3M (Ours)
72.13
68.85
60.66
71.43
57.14
54.29
77.08
79.17
70.83
55.17
55.17
48.28
95.45
77.27
86.36
73.33
68.21
63.08
Table 1: Main results under our online, session-by-session evaluation protocol on MemLens, across two execution models, five task categories, and three history lengths. All reported scores are accuracy percentages, shown without the % sign; boldface indicates the best scores among the compared methods.
Figure 3: Cross-session history pressure and active-index budget scaling. (a, b) With fixed C1 limits, history growth raises utilization and lowers task outcomes. (c, d) At fixed 128K history, active-index scaling reduces overflow repairs but yields non-monotonic outcomes; Raw-Node pointers remain fixed. Writer policy and reader budget are shared.
Figure 4: Case studies of multimodal evidence routing. C3M correctly identifies arched window frames (left), water bottles on the cooler’s right side (middle), and the McDonald’s sign above WEGO (right), while the baseline gives incorrect or inconclusive answers. Red boxes mark relevant evidence; green and red text indicate C3M and baseline answers, respectively.
Long-horizon agents can archive large histories, but future answers still incur retrieval, rereading, and context costs. When retained memory misses answer-relevant evidence, the system must return to larger portions of the raw history. We study budgeted evidence survival: before the query is known, which source evidence should be retained so that it remains recoverable and usable under a fixed retained source-evidence token budget? We instantiate this setting as Budgeted Pre-Query Retention, where memory is written during ingestion and later read without access to the full raw stream. We introduce EMBER, a learned retention policy that constructs a compact, source-backed evidence state. EMBER stores evidence capsules: verbatim source excerpts paired with retrieval keys and update metadata, preserving both grounding and read-time access. Post-query outcome feedback trains the writer to preserve evidence across the ingestion-retrieval-answer chain. On LongMemEval-RR, our LongMemEval-derived retained-evidence protocol, EMBER-14B reaches 0.3017 F1 at the 8192-token retained-evidence comparison point, compared with 0.1765 for the strongest non-EMBER budgeted baseline. Across retained source-evidence budgets, EMBER improves F1, Retain-Recall, and Read-Recall, indicating that long-horizon memory depends on retaining evidence within the budget rather than rereading larger histories.
Existing multimodal long-term memory agents use external memory to overcome the limited context available for long videos. However, most methods emphasize what to store rather than how stored memory should be retrieved. When retrieval becomes inaccurate or repeatedly fails to obtain useful evidence, existing agents lack mechanisms to diagnose failures from previous task trajectories and adapt future search strategies.We introduce Reflective Retrieval Memory (RRM), a reflective memory framework for long-horizon multimodal reasoning. RRM augments an entity-centric multimodal memory graph with reflective experience memory, which distills transferable procedural retrieval knowledge from historical task trajectories. Unlike episodic and semantic memories that preserve factual evidence from the current video, reflective experience memory captures reusable search strategies across tasks. RRM converts retrieved experiences into query-level guidance, while answer generation remains conditioned only on factual evidence newly retrieved from the current video. A lifecycle management mechanism further regulates experience memory through usage frequency, reuse feedback, and temporal decay, thereby reducing redundancy and noise. RRM consistently outperforms previous state-of-the-art approaches on M3-Bench-Robot, M3-Bench-Web, and Video-MME-Long, demonstrating the effectiveness of reflective retrieval memory for long-horizon multimodal reasoning.
Jingxiang Fan, Junbao Zhuo, Bochao Zou
University of Science and Technology Beijing 30 Xueyuan Road, Haidian District, Beijing 100083, China
Long-term memory is essential for LVLM agents to maintain consistency and integrate information across extended multimodal interactions. Existing agent memory systems, however, often reduce visual experiences into textual summaries or rely on static retrieve-then-reason pipelines, which are inefficient at query time and brittle when questions require image-text binding, temporal updates, or visual details. We propose Prospective Multimodal Memory Compilation, a framework that shifts part of the memory reasoning process from query time to memory consolidation time. Given accumulated multimodal interactions, a Questioner predicts future question candidates, a Planner compiles question-conditioned multimodal memory programs, and a Doubter verifies whether the planned evidence path can support the predicted answer. The verified question-program pairs form a structured question bank for efficient query-time routing and evidence retrieval. Experiments on multimodal long-term memory benchmarks show that our method improves answer quality and visual evidence recall while reducing query-time token and latency costs. Extensive ablations analyze the effects of self-feedback, dynamic planning, raw-image access, and question bank coverage.
Jingyu Sun, Yan Lin, Yuyang Xue +10
The University of Manchester · The University of Newcastle · The University of Edinburgh +5