Egocentric assistants must connect what users say with what they see across long interaction histories. We formalize this challenge as Spatially grounded Conversational Reasoning (SpaCR): cross-scene, recall-oriented, and counterfactual spatial queries that combine user-stated facts with geometric evidence. Direct vision-language models incur high inference costs and context limits as histories grow, while keyframe selection and retrieval can omit objects or evidence needed for complete recall. We propose Spatially grounded Conversational Memory (SpaC-MEM), an object-centric working memory that uses 3D reconstruction and segmentation to ground conversational information in persistent physical objects. It compresses multimodal histories while preserving spatial evidence and allowing object-specific facts to be updated through dialogue. We also introduce Ego-SpaCR, a benchmark comprising 620 ScanNet video sessions augmented with 95 task-oriented conversations and 3,100 evaluation queries. SpaC-MEM achieves the highest overall answer accuracy among the evaluated methods and improves object recall while requiring substantially fewer reference input tokens than native-video baselines. Removing 3D spatial information substantially degrades performance, highlighting the importance of preserving spatial and conversational evidence together.
Figures & tables
Figure 1: SpaCR involves complex spatial queries that require connecting user-stated facts to spatial evidence about the same physical objects across the interaction history. Here, identifying which of Dana’s belongings could fit in a box requires recalling ownership established through conversation and checking previously observed locations and object dimensions. SpaC-MEM links this conversational and spatial evidence through persistent objects in a unified working memory.
Figure 2: Example SpaCR queries by recall type (columns) and reasoning level (rows). Unlike VideoQA, each query requires user-stated facts from conversation alongside spatial evidence.
Figure 3: SpaC-MEM consolidates egocentric observations into persistent 3D object records, associates dialogue turns, and supplies the resulting working memory with visible objects to an LLM. The example combines updated intent and geometry for hypothetical viewpoint reasoning (C4).
Figure 4: Ego-SpaCR ’s pipeline for natural, verifiable dialogue over egocentric histories.
Statistic
Value
Statistic
Value
Task types
5
Grounded objects per conversation
49.8±24.3
Multi-session conversations
95
Object classes per conversation
27.4±9.8 (60 total)
Video sessions (clips)
620 (6.4 h)
Evaluation questions
3,100
Semantic facts per conversation
201.2±118.8
Questions per conversation
32.6±14.9
Sessions per conversation
6.5±2.7 [2, 14]
Complexity mix C1 / C2 / C3 / C4 (%)
6.0 / 24.3 / 45.9 / 23.8
Video per conversation (min)
4.1±1.9
Answer type count / object / boolean (%)
85.4 / 13.4 / 1.2
Table 1: Summary of 95 conversations spanning 620 video sessions and 3,100 evaluation questions in real ScanNet rooms (mean ± population s.d.; brackets give [min, max]).
Method
Model
C1
C2
C3
C4
All
Random
–
4.3
12.0
21.9
7.9
15.1
Conversation Only
GLM-5.3 Flash
41.7
39.0
22.0
23.2
27.6
In-Context Spatial Understanding
Native Video
Gemini 3.7 Flash
74.3
55.9
31.5
37.3
41.4
Qwen3.8 27B
64.2
46.3
11.3
8.5
22.3
Uniform
GLM-5.3 Flash
38.3
43.3
20.5
16.8
26.2
Table 2: Answer Accuracy across selected baselines and SpaC-MEM (%); bold and underlined entries are the best and second-best values per column.
Figure 5: Left: reference input tokens across methods and accuracy as a function of input frames before queries. SpaC-MEM achieves the highest accuracy while using substantially fewer reference input tokens. Right: object recall and answer accuracy across recall breadth. Colors and markers are consistent across panels; SpaC-MEM maintains the high recall and accuracy as breadth increases.
Figure 6: Oracle ablations between the original frame, oracle object box, and oracle geometric measurement (illustrated on the right panel). Ground-truth boxes yield only marginal gains, while oracle geometric measurements mainly improve C3 but not C4.
Figure 7: Ablations of SpaC-MEM : object-centric memory needs 4.7× fewer tokens than per-frame memory (a, b), removing spatial evidence lowers accuracy (c), and LLM-predicted dialogue association stays close to human-annotated GT (d).
Appendix figures & tables8 assets
Supplementary material from the paper’s appendix.
Appendix
Visual input
Conversation
Objects
Evaluation
Spatial reasoning
Benchmark
Egocentric video
Real scenes
Dialogue history
Object-specific user semantics
Cross-view object continuity
View grounding
Spatial relations
Geometric properties
Recall breadth
First-person perception
EPIC-KITCHENS ( Damen et al., 2018 )
✓
✓
×
×
×
◐
×
×
×
Ego4D ( Grauman et al., 2022 )
✓
✓
×
×
✓
◐
◐
◐
✓
Situated dialogue and spatial understanding
Appendix
Table 3: Benchmark coverage of visual input, conversation, object identity, and evaluation. ✓ : explicit support; ◐: related or limited support; × : not specified.
Task
Purpose
Example properties p
Example persona
Inventory walkthrough
Record what is in the space for later use.
Owner, provenance, history, and alias.
Roommate; new tenant.
Guided room tour
Explain the space so another person can use it later.
Function, history, owner, and alias.
Welcoming host; prospective tenant.
Lab or office setup
Prepare a shared work area without disturbing fixed items.
Constraint, responsibility, maintenance, and priority.
Lab member; organizer.
Workspace setup
Decide how to arrange a usable workspace.
Preference, intention, constraint, and function.
Remote worker; interior designer.
Cleaning and organizing
Decide what should be cleaned, kept, moved, or discarded.
Maintenance, status, decision, and planning.
Organizer; new tenant.
Appendix
Table 4: Five everyday tasks used to guide the conversational facts in each episode. Each episode keeps one task and a compatible persona across all selected clips; Appendix B.2 describes their role in conversation generation.
Task and conversation context
Task and persona: The continuing goal and the speaker’s role and manner. Example: Cleaning and organizing; a considerate organizer; continue in the study.
Prior context: Earlier dialogue, established object facts, and unfinished discussion. Example: The living room lamp was marked for cleaning; the walkthrough continues.
Observation windows
Time and view: A video interval and its representative frame. Example: Study window w1 , frames 120–180; representative frame 145.
Visibility changes: Objects in view, newly encountered, revisited, or leaving view. Example: Two lamps and a cabinet appear; a previously seen sofa remains visible.
Spatial context: Relations available in the window that can support object descriptions. Example: Cabinet near sofa, available in w1 after both objects have been observed.
Appendix
Table 5: Selected inputs to the Writer, with illustrative field values from one organizing task. Spatial context supports object references without disclosing the spatial outcomes being tested. Visual descriptions are checked independently before inclusion; query answers are never supplied.
Method
Model
C1
C2
C3
C4
All
Random
–
4.3 ±3.0
12.0 ±2.3
21.9 ±2.1
7.9 ±1.9
15.1 ±1.3
Conversation Only
GLM-5.3 Flash
41.7 ±7.0
39.0 ±3.5
22.0 ±2.2
23.2 ±3.0
27.6 ±1.6
In-Context Spatial Understanding
Native Video
Gemini 3.7 Flash
74.3 ±6.2
55.9 ±3.5
31.5 ±2.4
37.3 ±3.5
41.4 ±1.7
Qwen3.8 27B
64.2 ±6.8
46.3 ±3.6
11.3 ±1.6
8.5 ±2.0
22.3 ±1.5
Uniform
GLM-5.3 Flash
38.3 ±7.0
43.3 ±3.6
20.5 ±2.1
16.8 ±2.7
26.2 ±1.6
Appendix
Table 6: Answer Accuracy across selected baselines and SpaC-MEM (%) with Wilson 95% confidence interval half-widths in gray; bold and underlined entries are the best and second-best values per column.
Figure 8: Object recall for the ablations in Fig. 7 : SpaC-MEM keeps about 70% recall as history grows while per-frame memory degrades (a) at 4.7× the token cost (b), spatial evidence drives recall on spatial queries (c), and LLM-predicted dialogue association retains most of the recall of human-annotated GT association (d).
Figure 9: Answer accuracy and object recall as the number of objects in Wk grows. Solid and dashed lines use GPT-5.6 Luna and GLM-5.3 Flash, respectively. Curves are smoothed trends with 95% episode-bootstrap intervals; markers show means within 25-object bins. The horizontal axis covers 95% of queries.
0.25 m
0.5 m
1.0 m
Instance recall (%)
70.7
91.1
97.3
Semantic recall (%)
65.6
82.2
85.4
Instance precision (%)
52.0
67.0
71.6
Median centroid error of matched objects: 16.2 cm
Recall at 3D box IoU 0.25: 51.7%
Appendix
Table 7: Reconstructed object records compared with ScanNet annotations across 39 scenes and 864 annotated objects. Instance matching uses centroid distance; semantic recall also requires agreement on the ScanNet200 class.
Figure 10: Reconstruction recall by centroid-distance threshold (left) and annotated object size at 0.5 m (right). The bands show the per-scene interquartile range. Instance recall is class-agnostic; semantic recall also requires the correct ScanNet200 label.