cs.AIOct 4, 2026

Have I Scene This Before? Spatially Grounded Conversational Memory for Complex Queries in Egocentric Assistants

Authors: Jiazhou Liang, Liam Gallagher, Kiko Chen, David Guo, Armin Toroghi, Yifan Simon Liu, Scott Sanner

Organizations: University of Toronto

Abstract

Egocentric assistants must connect what users say with what they see across long interaction histories. We formalize this challenge as Spatially grounded Conversational Reasoning (SpaCR): cross-scene, recall-oriented, and counterfactual spatial queries that combine user-stated facts with geometric evidence. Direct vision-language models incur high inference costs and context limits as histories grow, while keyframe selection and retrieval can omit objects or evidence needed for complete recall. We propose Spatially grounded Conversational Memory (SpaC-MEM), an object-centric working memory that uses 3D reconstruction and segmentation to ground conversational information in persistent physical objects. It compresses multimodal histories while preserving spatial evidence and allowing object-specific facts to be updated through dialogue. We also introduce Ego-SpaCR, a benchmark comprising 620 ScanNet video sessions augmented with 95 task-oriented conversations and 3,100 evaluation queries. SpaC-MEM achieves the highest overall answer accuracy among the evaluated methods and improves object recall while requiring substantially fewer reference input tokens than native-video baselines. Removing 3D spatial information substantially degrades performance, highlighting the importance of preserving spatial and conversational evidence together.

Figures & tables

Appendix figures & tables8 assets

Supplementary material from the paper’s appendix.

Appendix

Explore similar work

CardsList
  1. Keep It in Mind: User Centric Continual Spatial Intelligence Reasoning in Egocentric Video Streams

    Jun 13, 2026Yun Wang, Junbin Xiao, Han Lyu +6Stable Spatial UnderstandingSpatial Memory

  2. Reconstructing the Right Episode: Evaluating Interleaved Conversational Memory Beyond Long Context

    Aug 26, 2026Zhexi Feng, Ruiyi Zhang, Yongbo Yang +1Conversational MemoryReasoning Benchmark

  3. CapMem: A Benchmark for Caption-Based Episodic Memory in Egocentric Video

    Sep 15, 2026Dingli Liang, Yiqiao Xie, Yukai Huang +6Long Video Question AnsweringEpisodic Memory