APM-Bench: Benchmarking Cross-session Persistent Memory for Egocentric Streaming Video Assistants
Authors: Jianguo Huang, Jinming Liu, Qiyao Wang, Liang Xu, Jianhang Li, Zhimian Wen, Mingda Li, Shule Lu, +4 more
Organizations: Shanghai Jiao Tong University · Eastern Institute of Technology, Ningbo · Shenzhen Institutes of Advanced Technology, Chinese Academy of Sciences · Zhongguancun Academy, Beijing, China · Dalian University of Technology · Beihang University · Hong Kong Polytechnic University
To serve as real-world personal assistants, streaming video models need persistent memory that retains past experiences for later use. Yet existing streaming benchmarks and methods often focus on individual continuous videos or short clips, overlooking that real-world interactions are often intermittent and require memory to persist across interruptions. To fill this gap, we introduce APM-Bench, which reformulates real-world streaming interaction as multi-session life trajectories. It contains 549 sessions, 104 trajectories, and 2,719 candidates, spanning both objective and open-ended questions. Each session is a video with fine-grained annotations, and sessions within a trajectory revolve around related activities. Models then use persistent memory to answer questions about past sessions and provide proactive responses while maintaining real-time interaction. This raises challenges: persistent memory must be storable, selectively retain information, be injected at the right time, and remain efficient. Moreover, finite storage may leave required evidence unavailable, so assistants should recognize missing evidence. Therefore, we systematically evaluate general video models under different memory protocols and diverse specialized streaming memory systems, and test whether models acknowledge insufficient evidence. Our evaluation reveals a clear utility--latency--storage trade-off: existing methods still struggle to simultaneously achieve reliable long-term recall, low overhead, and effective proactive assistance across sessions. APM-Bench provides a comprehensive testbed for developing and comparing persistent memory systems under realistic streaming conditions. We hope it encourages future work that jointly considers utility, latency, and storage toward more practical persistent memory for real-world streaming assistants.
Figures & tables
Figure 1: In a streaming setting, once an interaction ends, the model can no longer directly access its visual stream because real-world interactions are not replayed; the assistant therefore needs persistent memory to retain prior experience. Across sessions, the assistant updates and reuses persistent memory for cross-session understanding and proactive assistance while continuing real-time perception. The example shows repeated collaborative dessert-making across multiple sessions.
Figure 2: Utility-latency-storage trade-off across evaluated methods. Bubble size represents storage cost per hour. All general video models are evaluated under Raw Video as Memory .
Benchmark
RTP
RET
PRO
OE
OBJ
SE
MS
EA
StreamingBench ( Lin et al., 2024 )
✓
✓
✓
✗
✓
✗
✗
✗
OVO-Bench ( Li et al., 2025 )
✓
✓
✓
✗
✓
✗
✗
✗
PhoStream ( Lu et al., 2026 )
✓
✓
✓
✓
✗
✗
✗
✗
EgoStream ( Forte et al., 2026 )
✓
✓
✗
✗
✓
✓
✗
✗
EgoServe ( Sitong et al., 2026 )
✗
✗
✓
✓
✗
✗
✗
✗
StreamArena ( Zhang et al., 2026b )
✓
✓
✓
✓
✗
✗
✗
✗
Table 1: Comparison of streaming video benchmarks. Most prior benchmarks evaluate models on a single continuous video; APM-Bench evaluates interactions across sessions separated by interruptions, simulating intermittent real-world use and testing whether memory from earlier sessions remains useful. It also tests whether models recognize when required historical evidence is unavailable. RTP: real-time perception; RET: retrospective tasks; PRO: proactive response; OE: open-ended candidates; OBJ: objective candidates; SE: storage-efficiency evaluation; MS: multi-session evaluation across related activities; EA: Evidence Availability-Aware evaluation.
Figure 3: Examples of intra-session and inter-session streaming tasks. Each session follows its original continuous wall-clock timeline. Across the 12 tasks, we characterize six streaming temporal formulations: Cross-session understanding and Real-time Perception each follow a unified setup, while Adaptive Response tasks adopt four distinct patterns (with TPG and MPA sharing one). Notations: ∙ denotes user instructions, such as forward questions or reminder registrations; TPG and MPA trigger autonomously without explicit user prompts. ∙ denotes probes used to evaluate whether the model should remain SILENT or INTERVENE . Real-time Perception : evidence and query co-occur in the current session; Cross-session Understanding : evidence comes from prior sessions; and Adaptive Response : evidence comes from the current session for intra-session tasks (ERA and RCR), or from prior sessions for inter-session tasks (MPA, PRM, and TPG).
Figure 4: APM-Bench statistics. Left: distribution of 2,719 candidates across 12 tasks and three capability families. Middle: cumulative video duration across sessions in each trajectory. Right: wall-clock span from the first session start to the last session end, including inter-session gaps.
Figure 5: APM-Bench construction pipeline. In Stage 1 , EgoLife and HD-EPIC are organized into activity-related trajectories and segmented into sessions; real-world timestamps from the source metadata are rendered onto session videos that lack visible timestamps. In Stage 2 , candidates are generated from the session videos and progressively filtered and refined through choice-blind review, timestamp filtering, video-agent review, and human verification, yielding 2,719 candidates across 104 trajectories and 549 sessions. See Appendix D.1 for details.
Evaluation
Output Format
Metric
Cross-session Understanding
4-way MCQA
Accuracy
Real-time Perception
4-way MCQA
Accuracy
Evidence Availability-Aware
5-way MCQA
Accuracy
Adaptive Response
Open-ended
Gated LLM-Judge
Table 2: Evaluation metrics across capability families and the Evidence Availability-Aware setting.
Method
Streaming
Persistent
CS Und.
Perception
Adaptive
TTFT (s)
Storage Cost
Overall
w/o Memory
SimpleStream (Recent-4)
✗
✗
27.21
53.06
32.48
0.56
–
37.58
Raw Video as Memory
Seed-2.0-Lite ( ByteDance Seed, 2026 )
–
–
54.47
50.75
41.88
–
3.01 GiB
49.03
Gemini 3.6 Flash ( Google DeepMind, 2026 )
–
–
69.37
64.20
47.05
–
60.21
Qwen3.8-27B ( Qwen Team, 2026 )
–
–
50.10
48.84
38.31
30.72
45.75
Table 3: Task performance and efficiency of the evaluated memory methods. Streaming and Persistent indicate support for streaming input and persistent memory, respectively. CS Und.: Cross-session Understanding ; Perception: Real-time Perception ; Adaptive: Adaptive Response . Storage cost is reported per video hour, and Overall is the mean of the three capability scores.
Model / Method
Cross-session Und.
Real-time Perception
Adaptive Response
Overall
Evidence Availability-Aware
ER
EST
TR
ACR
CT
OCR
STU
ERA
RCR
MPA
PRM
TPG
EA
EU
w/o Memory
SimpleStream
27.52
28.52
25.57
56.44
35.90
73.49
46.41
50.89
30.22
26.54
32.11
22.64
37.58
3.85
95.38
Raw Video as Memory
Seed-2.0-Lite
54.13
63.09
46.18
57.43
43.59
53.61
48.37
41.97
65.63
24.45
51.56
25.80
49.03
64.62
66.15
Gemini 3.6 Flash
73.70
71.81
62.60
67.33
42.05
87.95
59.48
47.73
73.62
26.52
61.80
25.58
60.21
85.38
59.23
Table 4: Task-level results across the three capability families and Evidence Availability-Aware evaluation . EA: evidence available in accessible session. EU: evidence unavailable detection.
Appendix figures & tables17 assets
Supplementary material from the paper’s appendix.
Appendix
System
Visual-language backbone
SimpleStream
Qwen3-VL-8B-Instruct
HERMES
Qwen2.5-VL-7B-Instruct
ReKV
LLaVA-OneVision-Qwen2-7B
FluxMem
Qwen2.5-VL-7B-Instruct
Flash-VStream
Qwen2-VL-7B
StreamForest
Qwen2-7B
Appendix
Table 5: Visual-language backbones of specialized memory systems and SimpleStream.
Model / Method
ERA
RCR
MPA
PRM
TPG
Overall P/N/BA
w/o Memory
SimpleStream
65.89
50.81
54.53
50.86
50.39
19.86/92.81/56.34
Raw Video as Memory
Seed-2.0-Lite
62.10
73.63
53.94
69.56
58.15
80.66/46.65/63.66
Gemini 3.6 Flash
63.80
82.76
51.58
77.55
52.16
74.33/59.04/66.69
Qwen3-VL-8B
58.49
51.47
56.80
52.58
58.18
64.79/44.28/54.54
Appendix
Table 6: Adaptive decision accuracy (%). Each task score averages the correct INTERVENE rate on positive probes and the correct SILENT rate on negative probes. Overall pools probes across the five tasks before reporting the positive rate ( P ), negative rate ( N ), and their balanced average ( BA ).
Model / Method
ERA
RCR
MPA
PRM
TPG
Average
w/o Memory
SimpleStream
84.45 59.1%
60.18 50.8%
49.04 54.5%
63.81 50.9%
48.11 50.4%
61.12
Raw Video as Memory
Seed-2.0-Lite
92.91 45.4%
88.52 73.6%
46.45 53.9%
72.52 69.6%
44.73 58.2%
69.02
Gemini 3.6 Flash
88.33 54.2%
88.84 82.8%
52.58 51.6%
78.64 77.6%
52.31 52.2%
72.14
Qwen3-VL-8B
78.80 47.6%
57.46 51.5%
46.47 56.8%
54.05 52.6%
43.84 58.2%
56.12
Appendix
Table 7: Response quality after the Section 3.3 gate (%). Judged probes are averaged within each available polarity, then by candidate, using Eq. 3 . Superscripts give the balanced gate pass rate for each task: the mean of the positive and negative probe pass rates. Positive ERA probes also require the correct MCQA answer. Average is the mean of the five task scores.
Method
Mean RTF
HERMES
0.204
ReKV
0.089
FluxMem
0.016
Flash-VStream
0.010
StreamForest
0.020
OASIS
1.223
Appendix
Table 8: Offline replay real-time factor (RTF) for specialized memory methods. RTF is processing time to the first output token divided by the duration of the causal video prefix; each entry is the mean over evaluated probes. Values above 1 indicate that processing cannot keep pace with a live 1-FPS video stream under this replay setting. Lower is faster.
Model
Video memory
Text memory
Seed-2.0-Lite
34.11
11.25
Gemini 3.6 Flash
45.73
14.58
Appendix
Table 9: Mean end-to-end latency (seconds) for proprietary models, measured from request submission to the complete response. Video memory replays prior sessions; text memory supplies saved summaries and the current causal prefix.
Method
Mean peak
Maximum peak
HERMES
330.05 MiB
330.05 MiB
ReKV
19,202.80 MiB
58,185.20 MiB
FluxMem
167.31 MiB
189.96 MiB
Flash-VStream
565.20 MiB
566.25 MiB
StreamForest
3,374.53 MiB
3,867.47 MiB
OASIS
2,523.27 MiB
7,823.59 MiB
Appendix
Table 10: Peak persistent-memory size within each trajectory, after completed sessions. Mean and maximum are over 104 trajectories; units are shown in the cells.
Method
Count
Rate
Method
Count
Rate
HERMES
43
0.80%
StreamForest
1,941
36.14%
ReKV
516
9.61%
OASIS
2
0.04%
FluxMem
199
3.71%
Video-SALMONN S
792
14.75%
Flash-VStream
492
9.16%
VST
1,990
37.05%
Appendix
Table 11: Unrecoverable output-format failures for specialized memory methods. Each method is evaluated on 5,371 probes; rate is count divided by 5,371. Recoverable formatting errors are excluded.
Model
Video memory
Text memory
Seed-2.0-Lite
14 (0.26%)
4 (0.07%)
Gemini 3.6 Flash
4 (0.07%)
3 (0.06%)
Qwen3.8-27B
27 (0.50%)
81 (1.51%)
Qwen3-VL-8B
15 (0.28%)
9 (0.17%)
InternVL3.5-8B
98 (1.82%)
81 (1.51%)
VideoLLaMA3-7B
875 (16.29%)
816 (15.19%)
Appendix
Table 12: Unrecoverable output-format failures for general video models under video and text memory. Counts and percentages use 5,371 probes per setting. SimpleStream uses only the four most recent frames and is reported once.
Model / Method
Original Accuracy
Evidence Available Accuracy
Evidence Unavailable Detection
Balanced Accuracy
w/o Memory
SimpleStream
28.85
3.85
95.38
49.62
Raw Video as Memory
Seed-2.0-Lite
62.69
64.62
66.15
65.38
Gemini 3.6 Flash
76.15
85.38
59.23
72.31
Qwen3.8-27B
58.46
59.23
46.92
53.08
Appendix
Table 13: Evidence Availability-Aware accuracy (%) on 260 questions. Original Accuracy is the accuracy on the original four-option questions under each system’s main evaluation protocol, before restricting history. The restricted-history setting retains the two most recent completed sessions; Balanced Accuracy averages correct answer selection when evidence remains available and insufficient-evidence detection otherwise.
Figure 6: Trajectory organization on 366 candidates. Oracle supplies only necessary evidence; Raw Lifelong adds intervening source video. Adaptive Response uses the Section 3.3 Gated Judge score.
Figure 7: Scores on all 12 tasks for six general video models under video- and text-memory protocols. SimpleStream is shown in the same figure for comparison. All axes use a 0–100 scale.
Figure 8: Scores on all 12 tasks for eight specialized streaming-memory methods and SimpleStream. All axes use a 0–100 scale.
Figure 9: Activity terms in the 104 trajectory titles. Larger words occur in more trajectories.
Task
Candidates
Multi-evidence
Multi-session
Same-day
Cross-day
ER
327
33 (10.1%)
23 (7.0%)
53.5%
46.5%
EST
298
87 (29.2%)
24 (8.1%)
55.0%
45.0%
TR
262
262 (100%)
173 (66.0%)
54.6%
45.4%
MPA
116
48 (41.4%)
19 (16.4%)
59.5%
40.5%
TPG
132
73 (55.3%)
23 (17.4%)
77.3%
22.7%
Appendix
Table 14: Evidence composition for inter-session tasks with explicit evidence annotations. Multi-evidence requires at least two evidence items; multi-session evidence spans at least two sessions. Rates use the candidate count in each row. Cross-day means that at least one decisive item precedes the query day.
Task
1 session
2 sessions
3 sessions
≥ 4 sessions
ER
160
80
45
42
EST
160
78
37
23
TR
69
74
50
69
MPA
63
27
12
14
PRM
91
13
3
7
TPG
85
26
8
13
Appendix
Table 15: Inter-session candidates by the distance from the query session to the earliest decisive evidence session. Distance one means the immediately preceding session; the last column pools distances of four or more sessions.
Figure 10: Mean organized trajectory duration and real-world span by source. The span runs from the first session start to the last session end, including gaps; organized duration sums the retained session videos.
Figure 11: Human review consoles. Top: Human Verify & Refine shows video, time targets, task criteria, and editable fields for the 3,249 candidates retained after automated screening. Bottom: Agreement Check presents source evidence and candidate fields for the independent 300-question audit.
A central role of personal-agent memory is to turn stored information and prior interactions into future-oriented assistance. In daily use, useful cues come from what the agent observes and how the user interacts with the agent, and the agent must carry them forward from the current request to similar future tasks. Existing memory benchmarks usually test dialogue recall or task improvement in isolation, leaving the trajectory from streaming observations to later assistance largely untested. We introduce StreamMemBench, a streaming benchmark that constructs a two-step task sequence around each evidence anchor from EgoLife egocentric streams. The initial task tests evidence use, while the follow-up task tests whether feedback and interaction experience are reused. Four metrics diagnose evidence recall, initial evidence use, feedback incorporation, and follow-up reuse. Experiments with eight memory systems across two backbones show that current systems often fail to use observed evidence or turn feedback into reliable follow-up behavior, even when evidence is stored or feedback is incorporated locally. StreamMemBench is publicly available at https://github.com/landian60/StreamMemBench.
Continuous episodic memory is a core capability for autonomous agents operating in dynamic, real-world environments, yet current streaming video benchmarks provide limited tools for diagnosing what models remember and for how long. We introduce Egostream, a diagnostic benchmark for streaming episodic memory evaluation in egocentric vision. \egostream organizes 2,250 curated questions along seven cognitive dimensions: detail, spatial, temporal, event, social, causal, and prospective memory. We introduce the Answer Validity Window (AVW), which specifies the temporal span an answer remains valid as the observed scene evolves. This allows us to expand the questions into 8,528 recall-conditioned evaluations, enabling controlled testing from instant to ultra-long-term recall while separating genuine model forgetting from natural world-state changes. We rigorously establish baseline performance through a unified streaming MLLM framework that compares several state-of-the-art memory-management mechanisms, covering sliding windows, attention sinks, KV-cache pruning, merging, and offloading. Experiments within a unified Qwen3-VL backbone reveal that comparable aggregate accuracies mask starkly different memory profiles. For instance, token pruning preserves fine-grained details and temporal structure significantly better than token merging, while quantized offloading rescues ultra-long-term recall. Ultimately, all mechanisms operate well below real-time (>1s per frame), and top performing methods ceil at about 45% accuracy, exposing critical gaps in current architectures. Egostream provides the diagnostic testbed needed to close these gaps. Project website, news and updates at: https://saroo25.github.io/Egostream/
Rosario Forte, Giuseppe Lando, Antonino Furnari
Department of Mathematics and Computer Science University of Catania
Streaming video understanding models must answer queries at any moment during an ongoing stream, using only what they have observed so far and under fixed memory and computation budgets. Existing methods address this by adding memory banks, retrieval modules, or visual token compression to preserve long-range history. However, strong recent-window baselines show that indiscriminate history injection can dilute current-scene perception, suggesting that the key challenge is not whether to use memory, but how to allocate it selectively. We formulate this as budgeted online latent evidence allocation and propose \textbf{SelectStream}, a selective latent-memory framework that keeps the current observation directly visible to a frozen VLM while exposing historical information only through a compact, query-conditioned evidence budget. Three coordinated mechanisms govern when to write, what to preserve, and how to retrieve: surprise-driven adaptive windowing, priority-preserving consolidation, and query-conditioned graph reasoning over a fixed-capacity latent memory graph. Retrieved evidence is calibrated and injected as latent tokens for answer generation, without replaying frames or growing the context with stream length. Experimental results show that SelectStream achieves strong online streaming performance and preserves general video understanding, reaching 82.67% on StreamingBench, 67.03% on OVO-Bench, and 74.4% average accuracy on offline video benchmarks, while outperforming strong recent-window baselines and prior streaming memory methods.
Haonan Ge, Yiwei Wang, Hang Wu +1
University of California, Santa Barbara · University of California, Merced · The University of Queensland