cs.CVSep 4, 2026

Linguistic Trajectory Encoding for Efficient Long-Horizon Spatial Memory in Embodied Agents

Authors: Tianyidan Xie, Shenyi Wang, Qiang Tang, Mingjie Wang, Zhicheng Qiu, Xuanfu Li, Zhan Xu, Jian Yang, +2 more

Organizations: Nanjing University · University of British Columbia · Zhejiang Sci-Tech University · Huawei Technologies Co., Ltd. · Tianjin University

Abstract

Embodied agents performing long-horizon tasks require a memory representation in which the state transitions of dynamic objects remain queryable in natural language across hours-to-days observation horizons. Existing systems either drop fine-grained motion (clip-level video-language embeddings), keep it only as raw coordinates (geometric SLAM), or organise it around immediate task context (agent working memories). None of them gives the agent a per-object timeline whose state transitions are themselves queryable in language. Our key contribution is \textbf{Linguistic Trajectory Encoding} (LTE), which compresses dynamic object motion histories via a hybrid representation combining natural language descriptions, sparse spatial anchors, and visual anchors. LTE adapts compression to motion complexity by anchoring periods without reliable observations to the last seen location, while representing motion with geometric waypoints and linguistic descriptions to preserve accuracy. To evaluate these capabilities across extended time horizons, we construct the \textbf{Spatial Memory Benchmark} (SMB) from EgoLife multi-day recordings, targeting capabilities absent in existing benchmarks: semantic trajectory retrieval and long-horizon object retrieval. On SMB, the LTE-based system achieves 45.3%45.3\% success in semantic trajectory retrieval and 48.7%48.7\% in long-horizon object retrieval, outperforming structured-memory and VLM baselines (best prior: 31.9%31.9\% and 34.4%34.4\%). LTE achieves trajectory compression by factors of 8.7×8.7\times to 26.1×26.1\times with sub-second query latency on 2424,h video. On Ego4D natural-language queries, the system reaches 28.75%28.75\% / 55.10%55.10\% R@1/R@5, +15.80+15.80 / +31.30+31.30 pts over EgoVLPv2.

Figures & tables

Appendix figures & tables8 assets

Supplementary material from the paper’s appendix.

Appendix

Explore similar work

CardsList
  1. EmbodiedMemory-Bench: Benchmarking Embodied Memory for Long-Horizon Embodied Tasks

    Sep 23, 2026Lizhou Liang, Xinyu Zhong, Miao Pan +7Embodied AgentsInteraction History

  2. eMEM: A Hybrid Spatio-Temporal Memory System For Embodied Agents

    Jun 2, 2026A. Haroon Rasheed, Maria KabtoulEmbodied AgentsFactual Recall

  3. WorldLines: Benchmarking and Modeling Long-Horizon Stateful Embodied Agents

    Jun 17, 2026Yehang Zhang, Jianchong Su, Haojian Huang +7Long-Horizon AgentsLong-Horizon Task Planning