Linguistic Trajectory Encoding for Efficient Long-Horizon Spatial Memory in Embodied Agents
Organizations: Nanjing University · University of British Columbia · Zhejiang Sci-Tech University · Huawei Technologies Co., Ltd. · Tianjin University
Abstract
Embodied agents performing long-horizon tasks require a memory representation in which the state transitions of dynamic objects remain queryable in natural language across hours-to-days observation horizons. Existing systems either drop fine-grained motion (clip-level video-language embeddings), keep it only as raw coordinates (geometric SLAM), or organise it around immediate task context (agent working memories). None of them gives the agent a per-object timeline whose state transitions are themselves queryable in language. Our key contribution is \textbf{Linguistic Trajectory Encoding} (LTE), which compresses dynamic object motion histories via a hybrid representation combining natural language descriptions, sparse spatial anchors, and visual anchors. LTE adapts compression to motion complexity by anchoring periods without reliable observations to the last seen location, while representing motion with geometric waypoints and linguistic descriptions to preserve accuracy. To evaluate these capabilities across extended time horizons, we construct the \textbf{Spatial Memory Benchmark} (SMB) from EgoLife multi-day recordings, targeting capabilities absent in existing benchmarks: semantic trajectory retrieval and long-horizon object retrieval. On SMB, the LTE-based system achieves success in semantic trajectory retrieval and in long-horizon object retrieval, outperforming structured-memory and VLM baselines (best prior: and ). LTE achieves trajectory compression by factors of to with sub-second query latency on ,h video. On Ego4D natural-language queries, the system reaches / R@1/R@5, / pts over EgoVLPv2.
Figures & tables
| Method | IoU=0.3 | IoU=0.5 | ||
| R@1 | R@5 | R@1 | R@5 | |
| Supervised | ||||
| EgoVLPv2 ( Pramanick et al., 2023 ) | 12.95 | 23.80 | 7.91 | 16.11 |
| GroundNLQ ( Hou et al., 2023 ) | 27.20 | 54.42 | 18.91 | 39.98 |
| EgoVideo ( Pei et al., 2024 ) | 28.65 | 53.30 | 19.73 | 40.42 |
| OSGNet ( Feng et al., 2025 ) | 32.56 | 59.82 | 22.74 | 46.35 |
| Method | stAP | tAP | Succ. | Rec. |
| Supervised | ||||
| VQLoC | 0.22 | 0.31 | 55.9 | 47.1 |
| PRVQL | 0.27 | 0.35 | 57.9 | 47.9 |
| Zero-shot | ||||
| RELOCATE | 0.33 | 0.41 | 58.0 | 50.5 |
| Ours | 0.36 | 0.43 | 59.5 | 51.2 |
| Method | stAP | tAP | Succ. | Rec. |
| Supervised | ||||
| VQLoC | 0.22 | 0.31 | 55.9 | 47.1 |
| PRVQL | 0.27 | 0.35 | 57.9 | 47.9 |
| Zero-shot | ||||
| RELOCATE | 0.33 | 0.41 | 58.0 | 50.5 |
| Ours | 0.36 | 0.43 | 59.5 | 51.2 |
| Method | STR (%) | LOR (%) |
| VLM baselines (clip-scanning) | ||
| Q3VL-8B+GD | 21.5 | 25.1 |
| Q3VL-235B+GD | 31.9 | 34.4 |
| Structured-memory baselines | ||
| KFMem (3D-Mem-style) | 19.8 | 33.8 |
| VideoAgent ( Fan et al., 2024 ) | 24.7 | 30.5 |
| Dur. | #Obj. | Dense | Total | LTE | Compr. | Ours | Q3VL-235B+GD |
| (MB) | (MB) | (MB) | (s/q) | (s/q) | |||
| h | 18 | 71 | 45 | 8.2 | 0.15 | 8.2 | |
| h | 35 | 214 | 73 | 16.1 | 0.24 | 24.5 | |
| h | 52 | 428 | 98 | 23.5 | 0.32 | 49.1 | |
| h | 78 | 856 | 134 | 32.8 | 0.43 | 98.3 |
| Absolute (%) | vs. Ours | |||||||
| Configuration | NLQ | VQ2D | STR | LOR | NLQ | VQ2D | STR | LOR |
| Ours (full) | 55.10 | 59.5 | 45.3 | 48.7 | – | – | – | – |
| w/o text captions | 53.40 | 58.9 | 33.5 | 48.7 | ||||
| w/o spatial anchors | 54.30 | 57.1 | 41.2 | 48.7 | ||||
| w/o visual anchors | 54.70 | 53.4 | 43.9 | 48.7 | ||||
| w/o LTE (all removed) | 52.31 | 51.8 | 28.8 | 48.7 | ||||
Appendix figures & tables8 assets
Supplementary material from the paper’s appendix.
Appendix
| Lookback Window | Count | With Spatial Hint | Avg. Duration |
| hours | ( ) | – s | |
| hours | ( ) | – s | |
| hours | ( ) | – s | |
| hours | ( ) | – s | |
| Total | ( ) | – s |
| Component | Model | Role |
| Object detection | SAM3 (ViT-H) | Instance segmentation |
| Object tracking | SAM3 tracker | Cross-frame association |
| Point cloud | ViPE (ViT-L) | 3D reconstruction |
| VLM (scene) | Qwen3-VL-8B-Instruct | Room-level semantic labeling, captions |
| LLM (parsing) | Qwen3-8B | Query parsing, event extraction |
| Text embedding | Qwen3-Embedding-8B | Semantic similarity |
| Component | h video | h video | Parallelizable? |
| Offline memory construction | |||
| SAM3 detection + tracking | h | h | Yes (per-segment) |
| ViPE 3D reconstruction | h | h | Yes (independent) |
| VLM scene labeling | h | h | Yes (per-window) |
| VLM object captioning | h | h | After tracking |
| LTE construction + indexing | h | h | After captioning |
| Method | h | h | h | h | Avg. |
| Q3VL-8B+GD | |||||
| Q3VL-235B+GD | |||||
| Ours | |||||
| (Ours vs. Q3VL-235B) |
| Method | STR@ | STR@ | LOR@ | LOR@ |
| Q3VL-235B+GD | ||||
| Ours | ||||
| Query Type | Q3VL-235B+GD | Ours | |
| Single object queries | |||
| Two-object queries | |||
| With spatial hint only | |||
| With temporal hint only | |||
| With both hints | |||
| Overall |
| Configuration | STR | LOR | NLQ R@5 | VQ2D Succ. |
| Full LTE | ||||
| Text captions only | ||||
| Spatial anchors only | ||||
| Visual anchors only | ||||
| No LTE |
| Error Category | STR | LOR |
| Tracking failure (ID switch, lost track) | ||
| Caption ambiguity (imprecise description) | ||
| Spatial localization error (point cloud drift) | ||
| Query parsing error | ||
| Object occlusion (partial/full) |