AI agents require memory that preserves information across user-agent exchanges, user-to-user conversations, and group conversations with or without agent participation, while supporting updates as evidence changes or is removed. We present Stashbird, an agent memory system that links source episodes to derived memory state through explicit provenance. Stashbird organizes memory into episodic records, semantic relations, community summaries, and persisted graph state, with lifecycle operations for incremental updates and episode-level deletion. We evaluate question-answering accuracy and model-facing workload across four long-term memory benchmarks. On LoCoMo, Stashbird uses 76.4x fewer ingestion prompt tokens than Graphiti. Compared with reproduced Hindsight on the same benchmark, it uses 8.1x fewer retrieval prompt tokens, with accuracy 1.6 percentage points lower. It achieves higher accuracy than Hindsight on LongMemEval-S and GroupMemBench and comparable accuracy on EverMemBench.
Figures & tables
Figure 1: Overview of Stashbird. The purple region shows consolidation, where conversations are converted into episodes, summaries, semantic memories, and communities. The blue region shows persistent storage in the search index and graph database. The green region shows retrieval, where a query is decomposed, expanded through graph search, fused and reranked, and converted into the final context for the answer model.
View
Stored Unit
Update Behavior
Used For
Episodes
Conversation chunks
Append / delete
Evidence recall
Entities
Canonical entities
Merge / invalidate
Grounding
Relations
Entity relations
Replace / version
Reasoning
Preferences
User traits
Overwrite
Personalization
Communities
Entity clusters
Re-summarize
High-level context
Table 1: Structured memory views in Stashbird and their roles.