Never Look Back: Understanding Persistence in 3D Object Memory from Egocentric Videos
Organizations: Johns Hopkins University · Johns Hopkins University Applied Physics Laboratory
Abstract
As we move through the world and carry out everyday tasks, we encounter objects that may become relevant only later. We are capable of recalling where we left something or what was inside a container, even without knowing we would need it later. Here, we study how an embodied assistant can build a similar memory from egocentric videos, by observing a person's day-to-day activities. We present Ledger, a persistent 3D object memory that combines object locations, their histories, and contextual descriptions. It associates observations across the recording and retains objects after they leave the view, including those the person never touches. It clusters each object's observations by resting locations and records a move only after repeated evidence, reducing the effect of localization noise. Short descriptions preserve details such as an object's contents or supporting surface. It saves these records to later answer spatial questions without having to access the original images or video. Our memory raises HD-EPIC accuracy from 29.7% to 42.6%, UCS-Bench accuracy from 33.8% to 38.5% and localizes Ego4D objects with a 0.99 m median error on returned predictions. Our analyses identify complementary roles for temporal persistence, contextual descriptions, and retrieval. Our study on 100 stitched streams of multiple scenes each further exposes failures in both retrieval and construction. Per-scene construction partially recovers the performance lost across scene changes compared to that of single scene streams.
Figures & tables
| Memory content | Construction conditions | Query time | ||||||||
| Method | Handled objects | Untouched objects | Motion history | Path & places | RGB only | Scene changes | 20-min video | No frames at query | Language queries | 3D location |
| ConceptGraphs [ 18 ] | ✗ | ✓ | ✗ | ✗ | ✗ | ✗ | N.A. | ✓ | ✓ | ✓ |
| 3D-Mem [ 57 ] | ✗ | ✓ | ✗ | ✗ | ✗ | ✗ | N.A. | ✗ | ✓ | ✓ |
| ReMEmbR [ 2 ] | ✗ | ✗ | ✗ | ✓ | ✗ | N.A. | ✓ | ✓ | ✓ | |
| VideoAgent [ 14 ] | ✓ | ✗ | ✗ | ✓ | ✗ | ✗ | ✗ | ✓ | ✗ | |
| AMEGO [ 16 ] | ✓ | ✗ | ✓ | ✓ | ✗ | ✓ | ✓ | ✗ | ✗ | |
| Method | Video at answer | 3D Perception | Object Motion | Mean |
|---|---|---|---|---|
| LongVA [ 60 ] | yes | 32.9 | 22.7 | 27.8 |
| 2025 challenge, 1st place [ 42 ] | yes | 42.6 | 30.2 | 36.4 |
| EgoAdapt, 2026 [ 8 ] | yes | 64.9 | 61.6 | 63.3 |
| Human (sample) [ 35 ] | yes | 93.8 | 92.7 | 93.3 |
| Memories rebuilt on our frames, read by GPT-5.4 | ||||
| DirectMe, graph only [ 49 ] | no | 30.6 | 30.8 | 30.7 |
| Memory only | Frames + memory | |||||||
|---|---|---|---|---|---|---|---|---|
| Category | Blind | Frames | DM | Ledger | DM, own | DM, routed | Ledger | |
| Trajectory & Movement | 648 | 34.9 | 45.4 | 37.3 | 38.6 | 41.8 | 45.5 | 49.1 |
| Position & Orientation | 701 | 23.8 | 39.7 | 26.2 | 33.4 | 37.4 | 37.4 | 43.9 |
| Proximity & Reachability | 538 | 45.4 | 51.7 | 40.0 | 49.3 | 47.2 | 48.5 | 51.3 |
| Category & Quantity | 884 | 33.9 | 42.9 | 21.4 | 36.0 | 35.6 | 42.0 | 43.1 |
| Overall | 2771 | 33.8 | 44.4 | 30.0 | 38.5 | 39.8 | 42.9 | 46.3 |
| Memory | Graded / 164 | Median (m) | Mean (m) | m (%) | m (%) |
|---|---|---|---|---|---|
| ReMEmbR (wearer position) | 164 | 1.49 | 1.77 | 23 | 69 |
| DirectMe (graph only) | 164 | 1.91 | 1.98 | 18 | 52 |
| OSNOM (adapted) | 146 | 1.33 | 1.65 | 35 | 71 |
| No memory: clip centroid | 164 | 1.69 | 1.82 | 10 | 63 |
| Ledger , YOLO-World | 148 | 1.06 | 1.46 | 49 | 74 |
| Ledger , SAM 3 | 153 | 0.99 | 1.41 | 51 | 76 |
Appendix figures & tables40 assets
Supplementary material from the paper’s appendix.
Appendix
| Qwen3.5-9B | GPT-5.4 | |||||
| Category | Prototype | Blind | Memory | Blind | Memory | |
| 3D Perception | object location | 500 | 25.4 | 45.6 | 38.6 | 48.0 |
| contents retrieval | 200 | 16.0 | 28.5 | 16.5 | 33.5 | |
| fixture location | 500 | 16.2 | 52.6 | 13.8 | 56.0 | |
| fixture interaction counting | 300 | 36.0 | 40.3 | 40.0 | 41.3 | |
| Object Motion | movement itinerary | 500 | 11.8 | 26.2 | 34.8 | 38.6 |
| Answerer | Evidence | 3D Perception | Object Motion | Mean |
|---|---|---|---|---|
| Qwen3.5-9B | Blind | 23.4 [21.2,25.7] | 22.1 [19.0,25.1] | 22.8 [20.9,24.7] |
| Memory | 41.8 [39.3,44.4] | 25.4 [22.3,28.4] | 33.6 [31.6,35.5] | |
| GPT-5.4 | Blind | 27.2 [25.0,29.6] | 32.1 [28.9,35.4] | 29.7 [27.7,31.7] |
| Memory | 44.7 [42.0,47.3] | 40.5 [37.2,44.3] | 42.6 [40.5,44.8] |
| Stage | Parameter | Value |
|---|---|---|
| Sampling | interval / cap | 4 s / 1,200 frames |
| Rectification | pinhole resolution | px |
| Tagging | model / token cap / vocabulary cap | Qwen3.5-9B / 130 / 80 |
| Detection | model | YOLO-World v8x-worldv2 |
| input / confidence / top- | 1408 px / 0.05 / 40 | |
| 3D lifting | per-frame / per-video / per-tag budget | 12 / / |
| Model | Images to answerer | 3D Perception | Object Motion |
| Llama 3.2, text only [ 35 ] | no | 22.3 | 25.5 |
| Gemini 1.5 Pro, text only [ 35 ] | no | 21.5 | 27.7 |
| VideoLLaMA 2 [ 11 ] | yes | 25.7 | 28.5 |
| LongVA [ 60 ] | yes | 32.9 | 22.7 |
| LLaVA-Video [ 61 ] | yes | 27.3 | 18.9 |
| Gemini 1.5 Pro [ 44 ] | yes | 32.5 | 20.8 |
| Prototype | Video- LLaMA 2 | LLaVA- Video | Gemini 1.5 Pro | 2025 2nd | Xu et al. 2026 | GPT blind | Ours + Qwen | Ours + GPT | |
|---|---|---|---|---|---|---|---|---|---|
| Fixture interaction counting | 300 | 17.7 | 16.3 | 35.3 | 29.0 | 46.0 | 40.0 | 40.3 | 41.3 |
| Fixture location | 500 | 18.8 | 21.8 | 20.8 | 34.2 | 48.2 | 13.8 | 52.6 | 56.0 |
| Object location | 500 | 31.0 | 30.6 | 32.4 | 49.8 | 64.0 | 38.6 | 45.6 | 48.0 |
| Object contents retrieval | 200 | 35.5 | 40.5 | 41.5 | 50.5 | 58.5 | 16.5 | 28.5 | 33.5 |
| Movement itinerary | 500 | 11.0 | 9.8 | 18.0 | 14.2 | 49.0 | 34.8 | 26.2 | 38.6 |
| Movement counting | 200 | 44.0 | 20.0 | 13.0 | 44.5 | 51.0 | 31.0 | 35.5 | 45.0 |
| Prototype | Blind | DirectMe | ReMEmbR | OSNOM | AMEGO | Ours | |
|---|---|---|---|---|---|---|---|
| Fixture location | 500 | 14 | 26 | 21 | 44 | 13 | 56 |
| Object location | 500 | 39 | 37 | 39 | 42 | 40 | 48 |
| Object contents retrieval | 200 | 16 | 20 | 30 | 21 | 22 | 34 |
| Fixture interaction counting | 300 | 40 | 39 | 25 | 35 | 47 | 41 |
| Movement itinerary | 500 | 35 | 16 | 14 | 19 | 27 | 39 |
| Movement counting | 200 | 31 | 37 | 32 | 16 | 34 | 45 |
| Prototype | Full | Blind | Descriptions | Persistence | Triangulation | |
| 3D Perception | ||||||
| Fixture location | 447 | 56.4 | 12.8 | |||
| Object location | 500 | 48.0 | 38.6 | |||
| Object contents retrieval | 200 | 33.5 | 16.5 | |||
| Fixture interaction counting | 270 | 42.2 | 39.6 | |||
| Object Motion | ||||||
| Component removed | Without | With | Gain (points) |
|---|---|---|---|
| Persistence segmentation | 37.3 | 42.8 | [+3.1,+7.9] |
| Segment descriptions | 37.9 | 42.8 | [+2.9,+6.9] |
| Multi-view triangulation | 40.5 | 42.8 | [+0.2,+4.4] |
| 3D Perception | Object Motion | ||||||||
| Change from the final system | Mean | obj. loc. | contents | fixt. loc. | fixt. count | itinerary | mov. count | stationary | |
| Final system (GPT-5.4) | 42.6 | – | 48.0 | 33.5 | 56.0 | 41.3 | 38.6 | 45.0 | 38.0 |
| Answering model: GPT-5.6 | 45.0 | ||||||||
| no memory, GPT-5.4 | 29.7 | ||||||||
| no memory, GPT-5.6 | 27.2 | ||||||||
| Lift depth: FastVGGT | 44.7 | ||||||||
| Retrieval | Qwen3.5-9B | GPT-5.4 |
|---|---|---|
| Blind control | 19.5 | – |
| 3D centre distance | 29.2 | 36.6 |
| 2D area among candidates within 1 m | 33.9 | 36.6 |
| Union of both sets | 30.4 | 36.3 |
| 3D centre, triangulated positions | – | 40.4 |
| Evidence | Accuracy (%) |
|---|---|
| Predicted memory, triangulated positions | 40.4 |
| Ground-truth queried-object trajectory, positions only | 52.2 |
| Ground-truth trajectory and supporting-fixture names | 77.6 |
| Any of four single-candidate answers, retrieved objects | 65.5 |
| Any of four single-candidate answers, non-retrieved objects | 64.9 |
| Rule | Mean count | Exact | Within | Bias |
|---|---|---|---|---|
| Threshold, m | 4.9 | 13.5% | 36.0% | +2.3 |
| Threshold, m | 2.3 | 22.5% | 57.3% | |
| Persistence, m, | 2.2 | 28.1% | 67.4% |
| Detector | Memory | Poses | Median | Mean | m | m | CS |
|---|---|---|---|---|---|---|---|
| SAM 3 | Raw | EgoLoc | 0.877 | 1.399 | 53 | 77 | 99 |
| Raw | FastVGGT | 1.064 | 1.504 | 45 | 76 | 99 | |
| Trajectory | EgoLoc | 0.984 | 1.388 | 52 | 78 | 99 | |
| Trajectory | FastVGGT | 1.057 | 1.449 | 45 | 77 | 99 | |
| Consolidated | EgoLoc | 0.955 | 1.360 | 53 | 78 | 99 | |
| Consolidated | FastVGGT | 1.057 | 1.421 | 45 | 76 | 99 |
| Detector / level | Poses | median [95% CI] | 0.5 m (%) | Mean , all (m) | ||
|---|---|---|---|---|---|---|
| SAM 3 / raw | EgoLoc | 71 | 0.002 | 24 31 | 1.38 1.34 | |
| SAM 3 / raw | FastVGGT | 67 | 0.029 | 22 28 | 1.53 1.49 | |
| SAM 3 / trajectory | EgoLoc | 76 | 0.004 | 21 32 | 1.46 1.42 | |
| SAM 3 / trajectory | FastVGGT | 67 | 0.064 | 24 30 | 1.54 1.51 | |
| SAM 3 / tracker | EgoLoc | 70 | 0.003 | 27 37 | 1.36 1.32 | |
| SAM 3 / tracker | FastVGGT | 69 | 0.112 | 25 32 | 1.43 1.40 |
| Method | Succ (%) | Succ ∗ (%) | (m) | Angle | QwP (%) |
|---|---|---|---|---|---|
| Ego4D baseline [ 17 ] | 1.22 | 30.77 | 5.98 | 1.60 | 1.83 |
| Ego4D ∗ , EgoLoc poses [ 30 ] | 73.78 | 91.45 | 2.05 | 0.82 | 80.49 |
| EgoLoc [ 30 ] | 80.49 | 98.14 | 1.45 | 0.61 | 82.32 |
| EgoLoc-v1 [ 31 ] | 81.13 | 98.10 | 1.45 | 0.55 | 84.73 |
| EAGLE [ 4 ] | 84.77 | 98.54 | 1.18 | 0.42 | 85.68 |
| Embodied VideoAgent [ 15 ] | 85.37 | 92.72 | 1.86 | n/a | 92.07 |
| Memory | Graded | Succ | Succ ∗ | Median | Mean | m | m |
| AMEGO | N.C. | N.C. | N.C. | N.C. | N.C. | N.C. | N.C. |
| ReMEmbR (wearer position) | 164 | 100.0 | 100.0 | 1.49 | 1.77 | 23 | 69 |
| DirectMe (graph only) | 164 | 99.4 | 99.4 | 1.91 | 1.98 | 18 | 52 |
| OSNOM (adapted) | 146 | 89.0 | 100.0 | 1.33 | 1.65 | 35 | 71 |
| No memory: clip centroid | 164 | 100.0 | 100.0 | 1.69 | 1.82 | 10 | 63 |
| SAM 3 / raw / EgoLoc | 143 | 86.6 | 99.3 | 0.95 | 1.41 | 52 | 77 |
| Memory | Median | Mean | m (%) | m (%) |
|---|---|---|---|---|
| Ledger (SAM 3, trajectory, EgoLoc) | 1.00 | 1.44 | 50 | 76 |
| Ledger (SAM 3, trajectory, FastVGGT) | 1.05 | 1.49 | 44 | 74 |
| OSNOM (adapted) | 1.33 | 1.65 | 34 | 72 |
| ReMEmbR | 1.48 | 1.74 | 22 | 71 |
| DirectMe (graph only) | 1.88 | 1.93 | 17 | 54 |
| Clip centroid (no memory) | 1.69 | 1.81 | 9 | 64 |
| Detector | s/frame | Detections/clip | Box IoU | 3D distance (m) | Median (m) |
|---|---|---|---|---|---|
| SAM 3 | 4.4–11.1 | 1,281 | – | – | 0.984 |
| YOLO-World | 0.051 | 424 | 0.928 | 0.439 | 1.089 |
| YOLOE-26 | 0.047 | 488 | 0.950 | 0.433 | 1.318 |
| YOLOE-11 | 0.049 | 471 | 0.961 | 0.415 | 1.404 |
| Tag source | GPT judge | Qwen judge | Tags/clip |
|---|---|---|---|
| RAM++ | 53 | 39 | 53 |
| Gemma-3-4B | 76 | 64 | 297 |
| Qwen3.5-9B RAM++ | 84 | 68 | – |
| Canonical union | 86 | 70 | 112 |
| GPT vision, full | 87 | 76 | 177 |
| Gemma Qwen RAM++ | 87 | 75 | – |
| Text only | Frames + | ||||||||
|---|---|---|---|---|---|---|---|---|---|
| Subcategory | Blind | Ours | DM graph | DM blind | Frames | Ours | DM, own | DM, routed | |
| Quantity Change Tracking | 624 | 34.5 | 33.3 | 17.5 | 21.3 | 37.8 | 38.8 | 30.9 | 37.7 |
| My Trajectory | 465 | 35.9 | 41.3 | 38.3 | 43.2 | 45.8 | 48.8 | 41.5 | 44.1 |
| Relative Position & Orientation | 392 | 24.0 | 32.4 | 29.1 | 22.2 | 37.8 | 40.6 | 35.7 | 36.2 |
| Distance Comparison | 313 | 41.9 | 44.7 | 38.3 | 36.1 | 48.9 | 48.2 | 43.1 | 43.8 |
| Ego-centric Position | 309 | 23.6 | 34.6 | 22.7 | 23.6 | 42.1 | 48.2 | 39.5 | 38.8 |
| Dataset | Blind | Memory | 48 frames | Frames + memory | |
|---|---|---|---|---|---|
| EgoLife (189 recordings) | 2,569 | 33.6 | 38.4 | 44.1 | 45.8 |
| TeleEgo (22 recordings) | 203 | 37.4 | 39.4 | 47.3 | 52.2 |
| HourVideo (21 of 167 recordings) | 358 | 39.7 | 41.1 | 53.1 | 49.4 |
| All three | 3,130 | 34.5 | 38.8 | 45.4 | 46.6 |
| Evidence | All | Objects | Actions | People |
|---|---|---|---|---|
| Blind | 34.6 | 37.2 | 28.0 | 36.1 |
| Memory, text | 33.3 | 33.6 | 29.8 | 37.0 |
| Memory, text + counting aid (tracks, most visible together) | 26.9 | 25.0 | 25.0 | 34.5 |
| 48 uniform frames at 448 px | 37.9 | 33.9 | 41.1 | 45.4 |
| 48 frames chosen from memory keyframes | 37.4 | 33.6 | 36.9 | 49.6 |
| 24 frames at 896 px, half from the last 24 s, + 6 keyframes | 39.5 | 38.4 | 42.3 | 37.8 |
| All (3,244) | Same scene (1,959) | Scene change (1,285) | |||||||
|---|---|---|---|---|---|---|---|---|---|
| Evidence | single | stitched | single | stitched | single | stitched | |||
| None (blind) | 33.2 | 33.5 | 33.0 | 32.9 | 33.4 | 34.4 | |||
| 48 frames | 47.8 | 44.2 | 44.0 | 41.2 | 53.7 | 48.6 | |||
| DM graph | 30.1 | 29.6 | 29.7 | 29.7 | 30.7 | 29.6 | |||
| DM frames + graph (own) | 42.9 | 40.9 | 38.9 | 39.9 | 48.9 | 42.4 | |||
| Ledger memory | 39.0 | 36.4 | 38.4 | 37.9 | 39.8 | 34.2 | |||
| Subset | Blind | Frames | Ours mem. | Ours fr.+mem. | DM graph | DM fr.+graph | DM fr.+graph, routed | |
|---|---|---|---|---|---|---|---|---|
| All | 3244 | |||||||
| First recording | 1503 | |||||||
| After one earlier scene | 1558 | |||||||
| After two earlier scenes | 183 | |||||||
| Scene-change streams | 1285 | |||||||
| Same-scene streams | 1959 |
| Memory over the stitched stream | Accuracy | vs. single-scene build ( ) |
|---|---|---|
| Single-recording build (reference) | 39.8 | ( ) |
| Single-scene build over the stream | 34.2 | – |
| + retrieval restricted to the current scene, true cuts | 35.3 | (0.37) |
| + retrieval restricted, cuts from the place-partition detector | 35.0 | (0.34) |
| Per-scene construction, true cuts | 38.4 | ( ) |
| Per-scene construction, detected cuts | 37.0 | (0.05) |
| Prototype | Control | Event lines | [95% CI] | AMEGO |
|---|---|---|---|---|
| Object location | 48.8 | 50.8 | 39.8 | |
| Object contents retrieval | 30.5 | 35.0 | 22.0 | |
| Fixture location | 56.2 | 58.0 | 12.8 | |
| Fixture interaction counting | 41.3 | 40.7 | 47.0 | |
| Movement itinerary | 40.4 | 42.2 | 26.8 | |
| Movement counting | 42.5 | 43.0 | 34.5 |
| Memory | All | MyTraj | ObjTraj | EgoPos | RelPos | Dist | Reach | Recall | Count |
|---|---|---|---|---|---|---|---|---|---|
| 2,246 | 348 | 138 | 195 | 329 | 198 | 218 | 268 | 552 | |
| Full (objects + event lines) | 39.0 | 35.6 | 39.9 | 37.4 | 40.7 | 47.0 | 57.8 | 42.5 | 28.3 |
| event lines | 37.8 | 37.1 | 36.2 | 36.4 | 37.1 | 43.9 | 54.1 | 36.9 | 31.2 |
| object descriptions | 37.6 | 38.2 | 39.1 | 38.5 | 33.7 | 44.9 | 55.5 | 39.2 | 28.4 |
| 3D geometry | 38.6 | 37.6 | 41.3 | 24.1 | 38.9 | 48.0 | 56.9 | 46.3 | 29.3 |
| observation history | 38.4 | 39.9 | 42.0 | 29.7 | 33.7 | 49.5 | 56.9 | 39.9 | 30.4 |
| Record reader | All | MyTraj | ObjTraj | EgoPos | RelPos | Dist | Reach | Recall | Count |
|---|---|---|---|---|---|---|---|---|---|
| 2,777 | 423 | 156 | 270 | 409 | 288 | 257 | 307 | 667 | |
| ReMEmbR captions R1 | 36.0 | 38.5 | 37.8 | 29.6 | 32.3 | 41.7 | 55.3 | 32.9 | 30.6 |
| ReMEmbR captions R2 (= ReMEmbR) | 41.1 | 45.9 | 50.0 | 28.1 | 37.9 | 42.7 | 48.6 | 44.3 | 37.9 |
| Our event lines only R1 | 38.2 | 40.2 | 39.7 | 26.7 | 33.5 | 46.2 | 58.0 | 41.4 | 31.8 |
| Our objects + ReMEmbR captions R1 | 38.5 | 39.5 | 36.5 | 35.6 | 38.9 | 45.8 | 55.6 | 39.7 | 29.1 |
| Ledger (objects + event lines) R1 | 39.9 | 37.1 | 40.4 | 36.7 | 40.6 | 45.8 | 58.8 | 45.6 | 30.0 |
| Arm | All | MyTraj | ObjTraj | EgoPos | RelPos | Dist | Reach | Recall | Count |
|---|---|---|---|---|---|---|---|---|---|
| 3,300 | 511 | 191 | 305 | 480 | 318 | 308 | 374 | 813 | |
| Qwen3.5-9B, blind | 32.3 | 39.5 | 46.6 | 22.3 | 28.5 | 38.1 | 49.0 | 34.2 | 20.9 |
| Ledger + event lines, Qwen3.5-9B | 36.2 | 38.0 | 35.6 | 36.1 | 34.8 | 42.8 | 53.6 | 38.0 | 26.3 |
| DirectMe graph + keyframes, Qwen3-VL-8B | 28.0 | 33.3 | 33.5 | 24.9 | 27.1 | 36.5 | 46.8 | 26.2 | 15.5 |
| Ledger + event lines, GPT-5.4 | 39.5 | 36.2 | 38.7 | 37.0 | 39.8 | 45.9 | 58.1 | 44.7 | 30.6 |
| ReMEmbR, GPT-5.4 | 40.8 | 45.0 | 50.3 | 27.5 | 38.5 | 42.5 | 49.4 | 43.6 | 37.0 |
| Evidence | At its moment | Delayed | net of blind | |
|---|---|---|---|---|
| Blind | 32.4 | 37.9 | – | |
| 48 frames | 44.0 | 39.6 | ||
| DirectMe graph | 25.9 | 28.3 | ||
| ReMEmbR | 40.6 | 42.7 | ||
| Ledger , objects only | 34.1 | 32.8 | ||
| Ledger + event lines | 35.1 | 36.6 |
| Ledger | |||||||||
|---|---|---|---|---|---|---|---|---|---|
| Length | Videos | Questions | single-shot | agent | ReMEmbR | AMEGO | DirectMe | OSNOM | Blind |
| 5 | 45 | 284 | 48.6 | 35.6 | 29.6 | 35.2 | 33.8 | 38.4 | 38.7 |
| 5–10 | 22 | 266 | 48.5 | 39.1 | 32.7 | 31.6 | 27.1 | 32.7 | 30.5 |
| 10–20 | 37 | 769 | 45.1 | 33.2 | 27.8 | 31.3 | 29.6 | 31.1 | 28.9 |
| 20–40 | 34 | 710 | 42.4 | 30.1 | 23.4 | 31.1 | 30.3 | 32.7 | 28.0 |
| 40–75 | 12 | 371 | 41.8 | 37.2 | 29.9 | 31.3 | 26.4 | 30.5 | 27.0 |
| Ledger | ||||||||
|---|---|---|---|---|---|---|---|---|
| Length | Videos | Questions | single-shot | agent | ReMEmbR | DirectMe | Blind | 48 frames |
| 2 | 16 | 236 | 36.0 | 41.5 | 37.7 | 31.4 | 33.1 | 61.4 |
| 2–6 | 23 | 370 | 38.1 | 43.5 | 42.4 | 28.6 | 32.4 | 50.0 |
| 6–11 | 18 | 321 | 45.8 | 49.5 | 46.1 | 29.9 | 34.9 | 55.1 |
| 11–13 | 123 | 2,050 | 38.6 | 40.4 | 40.8 | 29.3 | 32.6 | 44.5 |
| 13–20 | 26 | 287 | 37.3 | 44.3 | 41.5 | 36.2 | 35.2 | 48.1 |
| Ledger | ||||||||
|---|---|---|---|---|---|---|---|---|
| Length | Clips | Queries | single-shot | agent | ReMEmbR | DirectMe | OSNOM | Clip centroid |
| 5 | 16 | 66 | 56.1 | 59.1 | 22.7 | 10.6 | 34.8 | 7.6 |
| 8 | 26 | 93 | 43.0 | 45.2 | 23.7 | 21.5 | 30.1 | 11.8 |
| 16 | 2 | 5 | 20.0 | 0.0 | 20.0 | 40.0 | 0.0 | 20.0 |
| All | 44 | 164 | 47.6 | 49.4 | 23.2 | 17.7 | 31.1 | 10.4 |
| Arm | HD-EPIC | UCS-Bench | UCS-Bench, EPIC-Kitchens only |
|---|---|---|---|
| Ledger , single-shot reader | (0.028) | (0.69) | (0.74) |
| Ledger | (0.12) | (0.82) | (0.72) |
| ReMEmbR | (0.099) | (0.63) | (0.17) |
| DirectMe | (0.044) | (0.71) | (0.89) |
| AMEGO (adapted) | (0.029) | – | – |
| OSNOM (adapted) | (0.016) | – | – |
| Stage | What is lost | Questions affected | Evidence |
|---|---|---|---|
| Sampling (one frame per 4 s) | objects seen only briefly; views for triangulation | all | 2 s sampling did not help a ten-video pilot (App. B.6 ) |
| Tagging | objects outside the per-video vocabulary never enter memory | any about those objects | the best tag set misses about one VQ3D query object in ten (Table 23 ) |
| 3D lifting | depth along the viewing ray | metric localization | the error lies almost entirely along the ray (App. B.7 ) |
| Association | one object split into several tracks | counting | the track count is exact for 11% of UCS-Bench count questions (App. D.2 ) |
| Record content | continuous rest intervals; hand events | stationary localization, fixture-interaction counting | rest-segment starts are near chance; hand tracklets score higher (App. B.8 , Table 10 ) |
| Construction | labels and metric scale at a scene cut; tractability with length | after a scene change; hour-long recordings | RQ5; 21 of 167 HourVideo memories built (App. D ) |
| Memory | Read at answer time | Acc. (%) | Input tok | Output tok | Calls |
|---|---|---|---|---|---|
| Blind | question and options only | 29.7 | 412 | 184 | 1.0 |
| ReMEmbR | captions, retrieval agent | 29.8 | 25,201 | 207 | 4.1 |
| OSNOM (adapted) | tracked object map | 30.2 | 8,762 | 111 | 1.0 |
| DirectMe | scene graph, its own prompt | 30.7 | 2,362 | 5 | 1.0 |
| Ledger , Qwen3.5-9B | object memory | 33.6 | 1,938 | 381 | 1.0 |
| AMEGO (adapted) | hand–object interaction log | 34.7 | 792 | 124 | 1.0 |
| Evidence | Read at answer time | Acc. (%) | Input tok | Output tok | Calls |
| Blind | question and options only | 33.8 | 452 | 149 | 1.0 |
| DirectMe | scene graph | 29.9 | 1,937 | 2 | 1.0 |
| Ledger | object memory | 38.5 | 2,044 | 211 | 1.0 |
| Frames | 48 frames | 44.4 | 9,615 | 50 | 1.0 |
| Frames, DM routed | frames; graph on spatial q. | 42.9 | 10,406 | 49 | 1.0 |
| Frames, DM own | frames and graph always | 39.8 | 11,377 | 48 | 1.0 |