Memory Agent Benchmarks

Momentum

16 papers in the last four weeks, up 167% on the four weeks before. 0.2% of all new papers.

Jul 13Week of Sep 28

Latest papers 89

All topics
CardsList
  1. Persistent Memory in Multi-Agent LLM Inference: What It Costs, What It Buys, and When You Can Tell

    Oct 6, 2026Hochan Son, Kyungdoe Han, Jaehan Koh +3LLM Inference EfficiencyMulti-Agent LLM Systems

  2. PACMI: Provenance-Aware Cascading Memory Invalidation for Long-Term LLM Agents

    Oct 5, 2026Yiqi Wang, Jiaqi Liu, Jiaqi Zhang +4Data ProvenanceMemory Agent Benchmarks

  3. RealCompanion: Benchmarking Human Understanding from Reasoning over Longitudinal Real-World Conversations

    Oct 1, 2026Arman Behnam, Sunglyoung Kim, Jiayi Yu +2Memory Agent BenchmarksUser Modeling

  4. TAGGRAPH: Tag-Augmented Graphs for Graph Retrieval of Agent Persistent Histories

    Sep 29, 2026Yu-Shu Chen, Yu-Jung Liang, Pengtao XieGraph-Based RetrievalMemory Agent Benchmarks

  5. APM-Bench: Benchmarking Cross-session Persistent Memory for Egocentric Streaming Video Assistants

    Sep 29, 2026Jianguo Huang, Jinming Liu, Qiyao Wang +9Streaming Video UnderstandingEgocentric Video QA

  6. RoutePrism: Tracing Construction Order Effects in Agent Memory

    Sep 28, 2026Dong Xu, Zhangfan Yang, Jiantao Wu +5Agent MemoryLLM Agent Memory

  7. TRACE: Governing Memory Validity in Evolving Multi-Agent Systems

    Sep 27, 2026Wenjun Xiong, Shengtao Zhang, Shangding Gu +5LLM Agent MemoryPersistent Memory for Language Models

  8. CoMemBench: Benchmarking Collaborative Memory Boundaries across Multi-Agent Workflow Topologies

    Sep 26, 2026Sen Zhao, Ruiqi Kong, Zuyu Zhang +5Multi-Agent OrchestrationMulti-Agent Collaboration

  9. EPGM: Execution Provenance for Budgeted Agent Memory Retrieval

    Sep 22, 2026Yiqi Wang, Jinqian Ju, Jiaqi Zhang +4Agent MemoryData Provenance

  10. MemCalib: Benchmarking and Optimizing Memory Use in LLM Agents

    Sep 21, 2026Ruike Cao, Fanyu Zhao, Fugen Yao +5LLM Agent MemoryMemory Agent Benchmarks

  11. ICM-Bench: Person-Level Identity Reasoning in Multimodal Agents with Long-Term Memory

    Sep 3, 2026Shidu Ren, Yunze Liu, Xing Liu +3Multimodal MemoryMemory Agent Benchmarks

  12. UTILMEM: Benchmarking Evidence Utilization in Long-Term Conversational Memory

    Aug 31, 2026Peijun Qing, Fobo Shi, Soroush VosoughiLLM Agent MemoryMemory Agent Benchmarks

  13. Governed Persistent Memory: Source-Bound State Semantics and Fail-Closed Release for Long-Horizon Agents

    Aug 12, 2026Guodong XuAgent MemoryLong-Horizon Agent Evaluation

  14. Total Recall at What Cost? Benchmarking the Serving Cost of Agentic Memory Systems

    Aug 12, 2026Natchanon Pollertlam, Witchayut KornsuwannawitLLM ServingLLM Agent Memory

  15. From Test-Time Scaling to Reusable Memory: Measuring Crystallization in Text-to-SQL

    Aug 7, 2026Jiaqian Wang, Yutao Qi, Wenjin Hou +2LLM MemoryText-to-SQL

  16. ContextWeave: A Real-World Workflow Benchmark

    Aug 5, 2026Bo Wang, Yuqian Yao, Enxi Wang +25Agent MemoryAgentic Workflows

  17. Caching for the Future: Scrub Jay Episodic Memory Principles for Agent Memory Systems

    Aug 5, 2026Kartikey Singh Bhandari, Aarya Wadhwani, Dhruv Kumar +1Temporal IREpisodic Memory